Calibration
GDP.pdf → MMMU-Pro
MMMU-Pro is estimated from GDP.pdf with a Hill curve fitted on 17 models measured on both: y = 0.7508 + (0.8784 − 0.7508)·x^6.00 / (0.21601^6.00 + x^6.00), R² = 0.80, cross-validated error 2.6 pp. It is used for 13 estimates.
| Estimated model | GDP.pdf | MMMU-Pro | Source |
|---|---|---|---|
| Claude Fable 5.1 (max with fallback) | 26.2% | 84.8% | estimated ± 2.6 pp, high confidence |
| GPT-6 Astra (medium) | 30.4% | 86.4% | estimated ± 2.6 pp, high confidence |
| Muse Spark 1.3 (max) | 26.6% | 85.0% | estimated ± 2.6 pp, high confidence |
| GLM-5.3-Flash | 15.4% | 76.6% | estimated ± 2.6 pp, high confidence |
| GLM-5.3 (max) | 11.2% | 75.3% | estimated ± 2.6 pp, high confidence |
| K2 Horizon 375B A23B | 7.4% | 75.1% | estimated ± 2.6 pp, medium confidence |
| Nemotron 3 Ultra | 5.0% | 75.1% | estimated ± 2.6 pp, medium confidence |
| Claude Sonnet 5.5 (max with fallback) | 25.8% | 84.6% | estimated ± 2.6 pp, high confidence |
| Gemini 4 Argon (high) | 21.8% | 81.6% | estimated ± 2.6 pp, high confidence |
| MiMo-V2.6-Pro | 19.2% | 79.3% | estimated ± 2.6 pp, high confidence |
| Grok 4.7 (xhigh) | 20.0% | 80.0% | estimated ± 2.6 pp, high confidence |
| GPT-6.1 Sol (xhigh) | 31.8% | 86.7% | estimated ± 2.6 pp, high confidence |
| GPT-6.1 Sol (high) | 32.0% | 86.7% | estimated ± 2.6 pp, high confidence |