benchgap
Calibration

AA-MMMU-Pro → MMMU-Pro w/ Python

MMMU-Pro w/ Python is estimated from AA-MMMU-Pro with a Michaelis–Menten curve fitted on 9 models measured on both: y = 2.0000·x / (1.15717 + x), R² = 0.89, cross-validated error 1.6 pp. It is used for 72 estimates.

Estimated modelAA-MMMU-ProMMMU-Pro w/ PythonSource
Apodex 1.179.2%81.3%estimated ± 1.6 pp, high confidence
Apodex 1.1 Mini79.2%81.3%estimated ± 1.6 pp, high confidence
Claude 3 Haiku30.8%42.0%estimated ± 1.6 pp, medium confidence
Claude 4.1 Opus Thinking67.9%74.0%estimated ± 1.6 pp, high confidence
Claude 4 Sonnet62.4%70.1%estimated ± 1.6 pp, medium confidence
Claude Opus 4.5 Thinking74.0%78.0%estimated ± 1.6 pp, high confidence
Claude Opus 4.6 (Adaptive)75.4%78.9%estimated ± 1.6 pp, high confidence
Claude Opus 4.776.4%79.5%estimated ± 1.6 pp, high confidence
Claude Opus 4.7 (Adaptive)78.8%81.0%estimated ± 1.6 pp, high confidence
Claude Opus 584.7%84.5%estimated ± 1.6 pp, medium confidence
Claude Opus 5.587.7%86.2%estimated ± 1.6 pp, medium confidence
Claude Sonnet 4.670.6%75.8%estimated ± 1.6 pp, high confidence
Claude Sonnet 577.3%80.1%estimated ± 1.6 pp, high confidence
DeepSeek V4.1 Flash77.0%79.9%estimated ± 1.6 pp, high confidence
Gemini 1.5 Pro55.0%64.4%estimated ± 1.6 pp, medium confidence
Gemini 2.5 Flash65.5%72.3%estimated ± 1.6 pp, high confidence
Gemini 2.5 Pro74.9%78.6%estimated ± 1.6 pp, high confidence
Gemini 3.5 Flash-Lite79.0%81.1%estimated ± 1.6 pp, high confidence
Gemini 3.6 Flash83.2%83.7%estimated ± 1.6 pp, high confidence
Gemini 3.7 Flash85.5%85.0%estimated ± 1.6 pp, medium confidence
Gemini 3.8 Flash85.6%85.0%estimated ± 1.6 pp, medium confidence
Gemini 3 Flash78.6%80.9%estimated ± 1.6 pp, high confidence
Gemma 3 27B48.0%58.6%estimated ± 1.6 pp, medium confidence
Gemma 4 E2B44.6%55.6%estimated ± 1.6 pp, medium confidence
Gemma 4 E4B51.4%61.5%estimated ± 1.6 pp, medium confidence
GLM-5V-Turbo72.8%77.2%estimated ± 1.6 pp, high confidence
GPT-4.161.2%69.2%estimated ± 1.6 pp, medium confidence
GPT-4.1 mini58.7%67.3%estimated ± 1.6 pp, medium confidence
GPT-4.1 nano40.1%51.5%estimated ± 1.6 pp, medium confidence
GPT-4o mini41.5%52.8%estimated ± 1.6 pp, medium confidence
GPT-5.175.5%79.0%estimated ± 1.6 pp, high confidence
GPT-5.1-Codex72.5%77.0%estimated ± 1.6 pp, high confidence
GPT-5.1-Codex-Max72.5%77.0%estimated ± 1.6 pp, high confidence
GPT-5.2-Codex76.3%79.5%estimated ± 1.6 pp, high confidence
GPT-5.3 Codex78.5%80.8%estimated ± 1.6 pp, high confidence
GPT-5 (high)74.2%78.1%estimated ± 1.6 pp, high confidence
GPT-5 (medium)74.3%78.2%estimated ± 1.6 pp, high confidence
GPT-6.1 Sol86.0%85.3%estimated ± 1.6 pp, medium confidence
GPT-6 Astra86.9%85.8%estimated ± 1.6 pp, medium confidence
GPT-6 Luna79.7%81.6%estimated ± 1.6 pp, high confidence
GPT-6 Sol82.9%83.5%estimated ± 1.6 pp, high confidence
Grok 468.8%74.6%estimated ± 1.6 pp, high confidence
Grok 4.1 Fast48.4%59.0%estimated ± 1.6 pp, medium confidence
Grok 4.1 Fast (Reasoning)63.3%70.7%estimated ± 1.6 pp, medium confidence
Grok 4.580.4%82.0%estimated ± 1.6 pp, high confidence
Grok 4 Fast (Reasoning)61.8%69.6%estimated ± 1.6 pp, medium confidence
LFM2.5-VL-1.6B-Extract26.5%37.3%estimated ± 1.6 pp, medium confidence
Ling 3.0 Flash VL79.0%81.1%estimated ± 1.6 pp, high confidence
Llama 4 Maverick62.1%69.8%estimated ± 1.6 pp, medium confidence
Llama 4 Scout52.9%62.7%estimated ± 1.6 pp, medium confidence
MiMo-V2.6-Flash73.1%77.4%estimated ± 1.6 pp, high confidence
MiMo-V2-Omni69.9%75.3%estimated ± 1.6 pp, high confidence
Mistral Large 355.7%65.0%estimated ± 1.6 pp, medium confidence
Mistral Large 476.4%79.5%estimated ± 1.6 pp, high confidence
Mistral Medium 353.0%62.8%estimated ± 1.6 pp, medium confidence
Mistral Medium 3.5 128B64.9%71.9%estimated ± 1.6 pp, medium confidence
Mistral Small 456.8%65.8%estimated ± 1.6 pp, medium confidence
Mistral Small 4 (Reasoning)56.8%65.8%estimated ± 1.6 pp, medium confidence
Nemotron 3 Nano Omni 30B A3B53.2%63.0%estimated ± 1.6 pp, medium confidence
Nova Pro44.3%55.4%estimated ± 1.6 pp, medium confidence
o370.1%75.5%estimated ± 1.6 pp, high confidence
Phi-4 Multimodal Instruct14.5%22.3%estimated ± 1.6 pp, medium confidence
Qwen3.5-122B-A10B75.0%78.7%estimated ± 1.6 pp, high confidence
Qwen3.5-27B75.0%78.7%estimated ± 1.6 pp, high confidence
Qwen3.5-35B-A3B72.7%77.2%estimated ± 1.6 pp, high confidence
Qwen3.5 397B (Reasoning)52.7%62.6%estimated ± 1.6 pp, medium confidence
Qwen3.8-27B76.3%79.5%estimated ± 1.6 pp, high confidence
Qwen3.8-Flash-Next79.8%81.6%estimated ± 1.6 pp, high confidence
Qwen3.8 Max Preview82.8%83.4%estimated ± 1.6 pp, high confidence
Qwen3-Omni-30B-A3B-Instruct55.5%64.8%estimated ± 1.6 pp, medium confidence
Qwen3-Omni-30B-A3B-Thinking60.2%68.4%estimated ± 1.6 pp, medium confidence
Step 3.7 Flash75.3%78.8%estimated ± 1.6 pp, high confidence