Calibration
Vals MMLU-Pro → LABBench2
LABBench2 is estimated from Vals MMLU-Pro with a Michaelis–Menten + offset curve fitted on 6 models measured on both: y = 0.0000 + 2.0000·x / (1.26627 + x), R² = 0.46, cross-validated error 1.8 pp. It is used for 34 estimates.
| Estimated model | Vals MMLU-Pro | LABBench2 | Source |
|---|---|---|---|
| Claude Haiku 4.5 | 78.7% | 76.7% | estimated ± 1.8 pp, low confidence |
| Claude Opus 4.7 | 89.9% | 83.0% | estimated ± 1.8 pp, low confidence |
| Gemini 3.1 Flash-Lite | 86.2% | 81.0% | estimated ± 1.8 pp, low confidence |
| Gemini 3.1 Pro | 91.0% | 83.6% | estimated ± 1.8 pp, low confidence |
| Gemini 3.5 Flash-Lite | 85.8% | 80.8% | estimated ± 1.8 pp, low confidence |
| Gemini 3 Flash | 88.6% | 82.3% | estimated ± 1.8 pp, low confidence |
| GLM-4.5 | 81.2% | 78.1% | estimated ± 1.8 pp, low confidence |
| GLM-4.6 | 82.2% | 78.7% | estimated ± 1.8 pp, low confidence |
| GLM-4.7 | 82.7% | 79.0% | estimated ± 1.8 pp, low confidence |
| GLM-5.1 | 86.9% | 81.4% | estimated ± 1.8 pp, low confidence |
| GLM-5.2 | 86.7% | 81.3% | estimated ± 1.8 pp, low confidence |
| GLM-5.3 | 86.8% | 81.3% | estimated ± 1.8 pp, low confidence |
| GLM-5.3-Flash | 86.1% | 80.9% | estimated ± 1.8 pp, low confidence |
| Grok 4.20 | 86.3% | 81.1% | estimated ± 1.8 pp, low confidence |
| Grok 4.3 | 85.8% | 80.8% | estimated ± 1.8 pp, low confidence |
| Inkling | 86.3% | 81.1% | estimated ± 1.8 pp, low confidence |
| Kimi K2.6 | 87.6% | 81.8% | estimated ± 1.8 pp, low confidence |
| Laguna M.1 | 68.8% | 70.4% | estimated ± 1.8 pp, low confidence |
| Laguna XS.2 | 69.1% | 70.6% | estimated ± 1.8 pp, low confidence |
| Ling 3.0 Flash | 82.0% | 78.6% | estimated ± 1.8 pp, low confidence |
| MiMo-V2.5 | 82.9% | 79.1% | estimated ± 1.8 pp, low confidence |
| MiMo-V2.5-Pro | 84.6% | 80.1% | estimated ± 1.8 pp, low confidence |
| MiniMax M2.7 | 80.4% | 77.7% | estimated ± 1.8 pp, low confidence |
| MiniMax M3 | 84.2% | 79.9% | estimated ± 1.8 pp, low confidence |
| Mistral Medium 3.5 128B | 75.3% | 74.6% | estimated ± 1.8 pp, low confidence |
| Muse Spark | 87.3% | 81.6% | estimated ± 1.8 pp, low confidence |
| Muse Spark 1.1 | 88.7% | 82.4% | estimated ± 1.8 pp, low confidence |
| Muse Spark 1.2 | 88.3% | 82.2% | estimated ± 1.8 pp, low confidence |
| Nemotron 3 Ultra | 85.8% | 80.8% | estimated ± 1.8 pp, low confidence |
| Qwen3.5 Flash | 84.1% | 79.8% | estimated ± 1.8 pp, low confidence |
| Qwen3.6 Plus | 87.7% | 81.8% | estimated ± 1.8 pp, low confidence |
| Qwen3.7 Max | 89.3% | 82.7% | estimated ± 1.8 pp, low confidence |
| Qwen3.8-27B | 84.3% | 79.9% | estimated ± 1.8 pp, low confidence |
| Qwen3.8 Max | 88.6% | 82.3% | estimated ± 1.8 pp, low confidence |