Calibration
SWE Multilingual → MMLU-ProX
MMLU-ProX is estimated from SWE Multilingual with a Michaelis–Menten + offset curve fitted on 7 models measured on both: y = 0.5008 + 2.0000·x / (3.57312 + x), R² = 0.64, cross-validated error 1.7 pp. It is used for 38 estimates.
| Estimated model | SWE Multilingual | MMLU-ProX | Source |
|---|---|---|---|
| Claude Fable 5.1 | 89.1% | 90.0% | estimated ± 1.7 pp, low confidence |
| Claude Mythos 5 | 92.2% | 91.1% | estimated ± 1.7 pp, low confidence |
| Claude Opus 4.8 | 84.4% | 88.3% | estimated ± 1.7 pp, low confidence |
| Claude Opus 5 | 89.5% | 90.1% | estimated ± 1.7 pp, low confidence |
| Claude Opus 5.5 | 93.9% | 91.7% | estimated ± 1.7 pp, low confidence |
| Claude Sonnet 5 | 78.3% | 86.0% | estimated ± 1.7 pp, medium confidence |
| Claude Sonnet 5.5 | 90.3% | 90.4% | estimated ± 1.7 pp, low confidence |
| Composer 2 | 73.7% | 84.3% | estimated ± 1.7 pp, medium confidence |
| Composer 2.5 | 79.8% | 86.6% | estimated ± 1.7 pp, low confidence |
| DeepSeek V4 Flash 0731 | 73.3% | 84.1% | estimated ± 1.7 pp, medium confidence |
| DeepSeek V4 Pro 0813 | 76.2% | 85.2% | estimated ± 1.7 pp, medium confidence |
| dots3-note Preview | 75.7% | 85.0% | estimated ± 1.7 pp, medium confidence |
| Granite 4.2 30B | 41.9% | 71.1% | estimated ± 1.7 pp, low confidence |
| Granite 4.2 8B | 30.8% | 65.9% | estimated ± 1.7 pp, low confidence |
| Grok 4.5 | 78.0% | 85.9% | estimated ± 1.7 pp, medium confidence |
| Hy4 preview | 82.9% | 87.7% | estimated ± 1.7 pp, low confidence |
| Kimi K2.6 | 76.7% | 85.4% | estimated ± 1.7 pp, medium confidence |
| Laguna M.1 | 63.1% | 80.1% | estimated ± 1.7 pp, low confidence |
| Laguna S 2.1 | 78.5% | 86.1% | estimated ± 1.7 pp, low confidence |
| Laguna XS.2 | 57.7% | 77.9% | estimated ± 1.7 pp, low confidence |
| Laguna XS 2.1 | 63.1% | 80.1% | estimated ± 1.7 pp, low confidence |
| Ling 3.0 Flash | 72.4% | 83.8% | estimated ± 1.7 pp, medium confidence |
| LLaDA2.2-flash | 25.0% | 63.2% | estimated ± 1.7 pp, low confidence |
| LongCat-Flash-Lite-Sparse | 59.3% | 78.6% | estimated ± 1.7 pp, low confidence |
| MiniMax M2.7 | 76.5% | 85.3% | estimated ± 1.7 pp, medium confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 36.5% | 68.6% | estimated ± 1.7 pp, low confidence |
| Ornith-1.0-35B | 69.3% | 82.6% | estimated ± 1.7 pp, medium confidence |
| Ornith-1.0-397B | 78.9% | 86.3% | estimated ± 1.7 pp, low confidence |
| Ornith-1.0-9B | 52.0% | 75.5% | estimated ± 1.7 pp, low confidence |
| Ornith-1.5-35B-A3B | 71.4% | 83.4% | estimated ± 1.7 pp, medium confidence |
| Ornith-1.5-397B | 79.6% | 86.5% | estimated ± 1.7 pp, low confidence |
| Ornith-1.5-9B | 54.4% | 76.5% | estimated ± 1.7 pp, low confidence |
| Qwen3.6-27B | 71.3% | 83.3% | estimated ± 1.7 pp, medium confidence |
| Qwen3.6-35B-A3B | 67.2% | 81.7% | estimated ± 1.7 pp, low confidence |
| Qwen3.8-Flash-Next | 81.0% | 87.0% | estimated ± 1.7 pp, low confidence |
| Qwen3.8-Omni-Flash | 80.5% | 86.8% | estimated ± 1.7 pp, low confidence |
| Beam | 78.0% | 85.9% | estimated ± 1.7 pp, medium confidence |
| SWE-1.7 | 77.8% | 85.8% | estimated ± 1.7 pp, medium confidence |