Calibration
HLE w/o tools → MMLU-Pro
MMLU-Pro is estimated from HLE w/o tools with a Michaelis–Menten curve fitted on 7 models measured on both: y = 0.8685·x / (0.00587 + x), R² = 0.69, cross-validated error 2.5 pp. It is used for 35 estimates.
| Estimated model | HLE w/o tools | MMLU-Pro | Source |
|---|---|---|---|
| Claude Fable 5.1 | 60.9% | 86.0% | estimated ± 2.5 pp, low confidence |
| Claude Haiku 5.5 | 45.9% | 85.8% | estimated ± 2.5 pp, low confidence |
| Claude Mythos 5 | 59.0% | 86.0% | estimated ± 2.5 pp, low confidence |
| Claude Opus 4.7 (Adaptive) | 46.9% | 85.8% | estimated ± 2.5 pp, low confidence |
| Claude Opus 4.8 | 49.8% | 85.8% | estimated ± 2.5 pp, low confidence |
| Claude Opus 5 | 56.3% | 86.0% | estimated ± 2.5 pp, low confidence |
| Claude Opus 5.5 | 64.4% | 86.1% | estimated ± 2.5 pp, low confidence |
| Claude Sonnet 5 | 43.2% | 85.7% | estimated ± 2.5 pp, low confidence |
| Claude Sonnet 5.5 | 56.9% | 86.0% | estimated ± 2.5 pp, low confidence |
| Gemini 3.1 Pro | 45.4% | 85.7% | estimated ± 2.5 pp, low confidence |
| GLM-5.2 | 40.5% | 85.6% | estimated ± 2.5 pp, low confidence |
| GPT-5.4 | 39.8% | 85.6% | estimated ± 2.5 pp, medium confidence |
| GPT-5.4 mini | 28.2% | 85.1% | estimated ± 2.5 pp, medium confidence |
| GPT-5.4 nano | 24.3% | 84.8% | estimated ± 2.5 pp, medium confidence |
| GPT-5.4 Pro | 42.7% | 85.7% | estimated ± 2.5 pp, low confidence |
| GPT-5.5 | 41.4% | 85.6% | estimated ± 2.5 pp, low confidence |
| GPT-5.5 Pro | 43.1% | 85.7% | estimated ± 2.5 pp, low confidence |
| Grok 4.20 | 31.6% | 85.3% | estimated ± 2.5 pp, medium confidence |
| Hy4 preview | 43.4% | 85.7% | estimated ± 2.5 pp, low confidence |
| Inkling | 30.0% | 85.2% | estimated ± 2.5 pp, medium confidence |
| Inkling-Small | 31.6% | 85.3% | estimated ± 2.5 pp, medium confidence |
| Kimi K3 | 43.5% | 85.7% | estimated ± 2.5 pp, low confidence |
| MiMo-V2.5-Pro | 34.0% | 85.4% | estimated ± 2.5 pp, medium confidence |
| Muse Spark | 42.8% | 85.7% | estimated ± 2.5 pp, low confidence |
| Muse Spark 1.1 | 52.2% | 85.9% | estimated ± 2.5 pp, low confidence |
| Ornith-1.5-35B-A3B | 25.6% | 84.9% | estimated ± 2.5 pp, medium confidence |
| Ornith-1.5-397B | 44.6% | 85.7% | estimated ± 2.5 pp, low confidence |
| Ornith-1.5-9B | 20.2% | 84.4% | estimated ± 2.5 pp, medium confidence |
| Pareto 26.9 | 49.0% | 85.8% | estimated ± 2.5 pp, low confidence |
| Qwen3.8-27B | 30.8% | 85.2% | estimated ± 2.5 pp, medium confidence |
| Qwen3.8-Flash-Next | 35.9% | 85.5% | estimated ± 2.5 pp, medium confidence |
| Qwen3.8 Max | 43.6% | 85.7% | estimated ± 2.5 pp, low confidence |
| Beam | 36.2% | 85.5% | estimated ± 2.5 pp, medium confidence |
| Sakana Fugu | 47.2% | 85.8% | estimated ± 2.5 pp, low confidence |
| Sakana Fugu-Ultra | 50.0% | 85.8% | estimated ± 2.5 pp, low confidence |