benchgap
Calibration

MMMU-Pro → MMMU-Pro w/ Python

MMMU-Pro w/ Python is estimated from MMMU-Pro with a Hill curve fitted on 9 models measured on both: y = 0.6079 + (1.2000 − 0.6079)·x^6.00 / (0.88881^6.00 + x^6.00), R² = 0.99, cross-validated error 0.6 pp. It is used for 33 estimates.

Estimated modelMMMU-ProMMMU-Pro w/ PythonSource
Claude Opus 4.570.6%72.7%estimated ± 0.6 pp, high confidence
Claude Opus 4.677.3%78.7%estimated ± 0.6 pp, high confidence
Command A+63.0%67.5%estimated ± 0.6 pp, medium confidence
dots3-note Preview79.1%80.4%estimated ± 0.6 pp, high confidence
Gemini 3.1 Pro83.9%85.3%estimated ± 0.6 pp, medium confidence
Gemini 3.5 Flash83.6%85.0%estimated ± 0.6 pp, medium confidence
Gemini 3 Pro81.0%82.4%estimated ± 0.6 pp, high confidence
Gemma 4 12B69.1%71.5%estimated ± 0.6 pp, high confidence
Gemma 4 26B A4B73.8%75.4%estimated ± 0.6 pp, high confidence
Gemma 4 31B76.9%78.3%estimated ± 0.6 pp, high confidence
GPT-5.279.5%80.8%estimated ± 0.6 pp, high confidence
Grok 4.2075.2%76.7%estimated ± 0.6 pp, high confidence
Grok 4.378.1%79.5%estimated ± 0.6 pp, high confidence
Inkling73.5%75.1%estimated ± 0.6 pp, high confidence
Inkling-Small74.0%75.6%estimated ± 0.6 pp, high confidence
Interfaze Beta71.1%73.1%estimated ± 0.6 pp, high confidence
Kimi K2.578.5%79.8%estimated ± 0.6 pp, high confidence
Kimi K2.5 (Reasoning)78.5%79.8%estimated ± 0.6 pp, high confidence
LFM2.5-VL-3B30.5%60.9%estimated ± 0.6 pp, medium confidence
MiMo-V2.577.9%79.3%estimated ± 0.6 pp, high confidence
MiniMax M378.1%79.5%estimated ± 0.6 pp, high confidence
Muse Glimmer 30B74.0%75.6%estimated ± 0.6 pp, high confidence
Muse Spark80.4%81.7%estimated ± 0.6 pp, high confidence
Pareto 26.978.0%79.4%estimated ± 0.6 pp, high confidence
Qwen3.5 397B79.0%80.3%estimated ± 0.6 pp, high confidence
Qwen3.6-27B75.8%77.2%estimated ± 0.6 pp, high confidence
Qwen3.6-35B-A3B75.3%76.8%estimated ± 0.6 pp, high confidence
Qwen3.6 Plus78.8%80.1%estimated ± 0.6 pp, high confidence
Qwen3.7 Plus79.0%80.3%estimated ± 0.6 pp, high confidence
Qwen3.8 Max82.3%83.7%estimated ± 0.6 pp, high confidence
Seed 2.1 Pro81.6%83.0%estimated ± 0.6 pp, high confidence
Seed 2.1 Turbo80.1%81.4%estimated ± 0.6 pp, high confidence
Step 5 Preview76.0%77.4%estimated ± 0.6 pp, high confidence