benchgap
Calibration

GDP.pdf → MMMU-Pro

MMMU-Pro is estimated from GDP.pdf with a Hill curve fitted on 17 models measured on both: y = 0.7508 + (0.8784 − 0.7508)·x^6.00 / (0.21601^6.00 + x^6.00), R² = 0.80, cross-validated error 2.6 pp. It is used for 13 estimates.

Estimated modelGDP.pdfMMMU-ProSource
Claude Fable 5.1 (max with fallback)26.2%84.8%estimated ± 2.6 pp, high confidence
GPT-6 Astra (medium)30.4%86.4%estimated ± 2.6 pp, high confidence
Muse Spark 1.3 (max)26.6%85.0%estimated ± 2.6 pp, high confidence
GLM-5.3-Flash15.4%76.6%estimated ± 2.6 pp, high confidence
GLM-5.3 (max)11.2%75.3%estimated ± 2.6 pp, high confidence
K2 Horizon 375B A23B7.4%75.1%estimated ± 2.6 pp, medium confidence
Nemotron 3 Ultra5.0%75.1%estimated ± 2.6 pp, medium confidence
Claude Sonnet 5.5 (max with fallback)25.8%84.6%estimated ± 2.6 pp, high confidence
Gemini 4 Argon (high)21.8%81.6%estimated ± 2.6 pp, high confidence
MiMo-V2.6-Pro19.2%79.3%estimated ± 2.6 pp, high confidence
Grok 4.7 (xhigh)20.0%80.0%estimated ± 2.6 pp, high confidence
GPT-6.1 Sol (xhigh)31.8%86.7%estimated ± 2.6 pp, high confidence
GPT-6.1 Sol (high)32.0%86.7%estimated ± 2.6 pp, high confidence