benchgap
Calibration

MMMU-Pro → CharXiv w/o tools

CharXiv w/o tools is estimated from MMMU-Pro with a linear curve fitted on 5 models measured on both: y = 1.0831·x + -0.0223, R² = 0.94, cross-validated error 1.8 pp. It is used for 27 estimates.

Estimated modelMMMU-ProCharXiv w/o toolsSource
Claude Opus 4.677.3%81.5%estimated ± 1.8 pp, medium confidence
Command A+63.0%66.0%estimated ± 1.8 pp, low confidence
Gemini 3.1 Pro83.9%88.6%estimated ± 1.8 pp, low confidence
Gemini 3.5 Flash83.6%88.3%estimated ± 1.8 pp, low confidence
Gemma 4 26B A4B73.8%77.7%estimated ± 1.8 pp, medium confidence
Gemma 4 31B76.9%81.1%estimated ± 1.8 pp, medium confidence
GPT-5.481.2%85.7%estimated ± 1.8 pp, medium confidence
GPT-5.4 mini76.6%80.7%estimated ± 1.8 pp, medium confidence
GPT-5.4 nano66.1%69.4%estimated ± 1.8 pp, low confidence
GPT-5.581.2%85.7%estimated ± 1.8 pp, medium confidence
GPT-5.6 Luna78.4%82.7%estimated ± 1.8 pp, medium confidence
GPT-5.6 Sol83.0%87.7%estimated ± 1.8 pp, low confidence
GPT-5.6 Terra80.7%85.2%estimated ± 1.8 pp, medium confidence
Grok 4.2075.2%79.2%estimated ± 1.8 pp, medium confidence
Grok 4.378.1%82.4%estimated ± 1.8 pp, medium confidence
Interfaze Beta71.1%74.8%estimated ± 1.8 pp, low confidence
Kimi K2.578.5%82.8%estimated ± 1.8 pp, medium confidence
Kimi K2.5 (Reasoning)78.5%82.8%estimated ± 1.8 pp, medium confidence
LFM2.5-VL-3B30.5%30.8%estimated ± 1.8 pp, low confidence
MiMo-V2.577.9%82.1%estimated ± 1.8 pp, medium confidence
MiniMax M378.1%82.4%estimated ± 1.8 pp, medium confidence
Muse Glimmer 30B74.0%77.9%estimated ± 1.8 pp, medium confidence
Muse Spark80.4%84.9%estimated ± 1.8 pp, medium confidence
Pareto 26.978.0%82.3%estimated ± 1.8 pp, medium confidence
Qwen3.6-27B75.8%79.9%estimated ± 1.8 pp, medium confidence
Qwen3.6-35B-A3B75.3%79.3%estimated ± 1.8 pp, medium confidence
Step 5 Preview76.0%80.1%estimated ± 1.8 pp, medium confidence