benchgap
Calibration

AA-MMMU-Pro → MMMU

MMMU is estimated from AA-MMMU-Pro with a Hill curve fitted on 8 models measured on both: y = 0.6826 + (1.2000 − 0.6826)·x^5.83 / (0.87835^5.83 + x^5.83), R² = 0.98, cross-validated error 0.9 pp. It is used for 96 estimates.

Estimated modelAA-MMMU-ProMMMUSource
Apodex 1.179.2%86.6%estimated ± 0.9 pp, medium confidence
Apodex 1.1 Mini79.2%86.6%estimated ± 0.9 pp, medium confidence
Claude 3 Haiku30.8%68.4%estimated ± 0.9 pp, medium confidence
Claude 4.1 Opus Thinking67.9%77.7%estimated ± 0.9 pp, high confidence
Claude 4 Sonnet62.4%74.5%estimated ± 0.9 pp, high confidence
Claude Opus 4.571.2%80.0%estimated ± 0.9 pp, high confidence
Claude Opus 4.5 Thinking74.0%82.2%estimated ± 0.9 pp, high confidence
Claude Opus 4.672.5%81.0%estimated ± 0.9 pp, high confidence
Claude Opus 4.6 (Adaptive)75.4%83.3%estimated ± 0.9 pp, high confidence
Claude Opus 4.776.4%84.2%estimated ± 0.9 pp, high confidence
Claude Opus 4.7 (Adaptive)78.8%86.2%estimated ± 0.9 pp, medium confidence
Claude Opus 584.7%91.4%estimated ± 0.9 pp, medium confidence
Claude Opus 5.587.7%94.0%estimated ± 0.9 pp, medium confidence
Claude Sonnet 4.670.6%79.6%estimated ± 0.9 pp, high confidence
Claude Sonnet 577.3%84.9%estimated ± 0.9 pp, high confidence
DeepSeek V4.1 Flash77.0%84.7%estimated ± 0.9 pp, high confidence
Gemini 1.5 Pro55.0%71.4%estimated ± 0.9 pp, high confidence
Gemini 2.5 Flash65.5%76.2%estimated ± 0.9 pp, high confidence
Gemini 2.5 Pro74.9%82.9%estimated ± 0.9 pp, high confidence
Gemini 3.1 Pro82.4%89.4%estimated ± 0.9 pp, medium confidence
Gemini 3.5 Flash84.3%91.0%estimated ± 0.9 pp, medium confidence
Gemini 3.5 Flash-Lite79.0%86.4%estimated ± 0.9 pp, medium confidence
Gemini 3.6 Flash83.2%90.1%estimated ± 0.9 pp, medium confidence
Gemini 3.7 Flash85.5%92.1%estimated ± 0.9 pp, medium confidence
Gemini 3.8 Flash85.6%92.2%estimated ± 0.9 pp, medium confidence
Gemini 3 Flash78.6%86.0%estimated ± 0.9 pp, medium confidence
Gemini 3 Pro80.2%87.4%estimated ± 0.9 pp, medium confidence
Gemma 3 27B48.0%69.7%estimated ± 0.9 pp, medium confidence
Gemma 4 12B69.7%78.9%estimated ± 0.9 pp, high confidence
Gemma 4 26B A4B69.2%78.6%estimated ± 0.9 pp, high confidence
Gemma 4 31B73.4%81.7%estimated ± 0.9 pp, high confidence
Gemma 4 E2B44.6%69.2%estimated ± 0.9 pp, medium confidence
Gemma 4 E4B51.4%70.4%estimated ± 0.9 pp, medium confidence
GLM-5V-Turbo72.8%81.2%estimated ± 0.9 pp, high confidence
GPT-4.161.2%73.9%estimated ± 0.9 pp, high confidence
GPT-4.1 mini58.7%72.8%estimated ± 0.9 pp, high confidence
GPT-4.1 nano40.1%68.8%estimated ± 0.9 pp, medium confidence
GPT-4o mini41.5%68.9%estimated ± 0.9 pp, medium confidence
GPT-5.175.5%83.4%estimated ± 0.9 pp, high confidence
GPT-5.1-Codex72.5%81.0%estimated ± 0.9 pp, high confidence
GPT-5.1-Codex-Max72.5%81.0%estimated ± 0.9 pp, high confidence
GPT-5.2-Codex76.3%84.1%estimated ± 0.9 pp, high confidence
GPT-5.3 Codex78.5%85.9%estimated ± 0.9 pp, medium confidence
GPT-5.478.4%85.9%estimated ± 0.9 pp, medium confidence
GPT-5.4 mini73.3%81.6%estimated ± 0.9 pp, high confidence
GPT-5.4 nano65.4%76.1%estimated ± 0.9 pp, high confidence
GPT-5.579.9%87.2%estimated ± 0.9 pp, medium confidence
GPT-5.6 Luna78.6%86.0%estimated ± 0.9 pp, medium confidence
GPT-5.6 Sol83.4%90.3%estimated ± 0.9 pp, medium confidence
GPT-5.6 Terra80.7%87.9%estimated ± 0.9 pp, medium confidence
GPT-5 (high)74.2%82.3%estimated ± 0.9 pp, high confidence
GPT-5 (medium)74.3%82.4%estimated ± 0.9 pp, high confidence
GPT-6.1 Sol86.0%92.5%estimated ± 0.9 pp, medium confidence
GPT-6 Astra86.9%93.3%estimated ± 0.9 pp, medium confidence
GPT-6 Luna79.7%87.0%estimated ± 0.9 pp, medium confidence
GPT-6 Sol82.9%89.8%estimated ± 0.9 pp, medium confidence
Grok 468.8%78.3%estimated ± 0.9 pp, high confidence
Grok 4.1 Fast48.4%69.8%estimated ± 0.9 pp, medium confidence
Grok 4.1 Fast (Reasoning)63.3%74.9%estimated ± 0.9 pp, high confidence
Grok 4.378.1%85.6%estimated ± 0.9 pp, medium confidence
Grok 4.580.4%87.6%estimated ± 0.9 pp, medium confidence
Grok 4 Fast (Reasoning)61.8%74.2%estimated ± 0.9 pp, high confidence
Inkling73.5%81.8%estimated ± 0.9 pp, high confidence
Inkling-Small74.0%82.2%estimated ± 0.9 pp, high confidence
Kimi K2.679.4%86.7%estimated ± 0.9 pp, medium confidence
Kimi K2.575.4%83.3%estimated ± 0.9 pp, high confidence
Kimi K2.5 (Reasoning)75.4%83.3%estimated ± 0.9 pp, high confidence
Kimi K380.5%87.7%estimated ± 0.9 pp, medium confidence
LFM2.5-VL-1.6B-Extract26.5%68.3%estimated ± 0.9 pp, medium confidence
Ling 3.0 Flash VL79.0%86.4%estimated ± 0.9 pp, medium confidence
Llama 4 Maverick62.1%74.3%estimated ± 0.9 pp, high confidence
Llama 4 Scout52.9%70.8%estimated ± 0.9 pp, medium confidence
MiMo-V2.6-Flash73.1%81.5%estimated ± 0.9 pp, high confidence
MiMo-V2-Omni69.9%79.1%estimated ± 0.9 pp, high confidence
MiniMax M378.6%86.0%estimated ± 0.9 pp, medium confidence
Mistral Large 355.7%71.7%estimated ± 0.9 pp, high confidence
Mistral Large 476.4%84.2%estimated ± 0.9 pp, high confidence
Mistral Medium 353.0%70.8%estimated ± 0.9 pp, medium confidence
Mistral Medium 3.5 128B64.9%75.8%estimated ± 0.9 pp, high confidence
Mistral Small 456.8%72.0%estimated ± 0.9 pp, high confidence
Mistral Small 4 (Reasoning)56.8%72.0%estimated ± 0.9 pp, high confidence
Muse Glimmer 30B74.3%82.4%estimated ± 0.9 pp, high confidence
Muse Spark80.5%87.7%estimated ± 0.9 pp, medium confidence
Nova Pro44.3%69.2%estimated ± 0.9 pp, medium confidence
o370.1%79.2%estimated ± 0.9 pp, high confidence
Phi-4 Multimodal Instruct14.5%68.3%estimated ± 0.9 pp, medium confidence
Qwen3.5 397B52.7%70.8%estimated ± 0.9 pp, medium confidence
Qwen3.5 397B (Reasoning)52.7%70.8%estimated ± 0.9 pp, medium confidence
Qwen3.7 Plus80.5%87.7%estimated ± 0.9 pp, medium confidence
Qwen3.8-27B76.3%84.1%estimated ± 0.9 pp, high confidence
Qwen3.8-Flash-Next79.8%87.1%estimated ± 0.9 pp, medium confidence
Qwen3.8 Max Preview82.8%89.7%estimated ± 0.9 pp, medium confidence
Qwen3-Omni-30B-A3B-Instruct55.5%71.6%estimated ± 0.9 pp, high confidence
Qwen3-Omni-30B-A3B-Thinking60.2%73.4%estimated ± 0.9 pp, high confidence
Step 3.7 Flash75.3%83.2%estimated ± 0.9 pp, high confidence
Step 5 Preview76.4%84.2%estimated ± 0.9 pp, high confidence