benchgap
Vision & documents

MMMU-Pro w/ Python leaderboard

As of 2026-10-07, the highest measured score on MMMU-Pro w/ Python is 84.6% by GPT-5.6 Sol. 105 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.586.2%estimated ± 1.6 pp, medium confidence
2GPT-6 Astra85.8%estimated ± 1.6 pp, medium confidence
3Gemini 3.1 Pro85.3%estimated ± 0.6 pp, medium confidence
4GPT-6.1 Sol85.3%estimated ± 1.6 pp, medium confidence
5Gemini 3.8 Flash85.0%estimated ± 1.6 pp, medium confidence
6Gemini 3.5 Flash85.0%estimated ± 0.6 pp, medium confidence
7Gemini 3.7 Flash85.0%estimated ± 1.6 pp, medium confidence
8GPT-5.6 Sol84.6%measured
9Claude Opus 584.5%estimated ± 1.6 pp, medium confidence
10Qwen3.8 Max83.7%estimated ± 0.6 pp, high confidence
11Gemini 3.6 Flash83.7%estimated ± 1.6 pp, high confidence
12GPT-6 Sol83.5%estimated ± 1.6 pp, high confidence
13Qwen3.8 Max Preview83.4%estimated ± 1.6 pp, high confidence
14Kimi K383.4%measured
15GPT-5.583.2%measured
16Seed 2.1 Pro83.0%estimated ± 0.6 pp, high confidence
17Gemini 3 Pro82.4%estimated ± 0.6 pp, high confidence
18GPT-5.482.1%measured
19GPT-5.6 Terra82.0%measured
20Grok 4.582.0%estimated ± 1.6 pp, high confidence
21Muse Spark81.7%estimated ± 0.6 pp, high confidence
22Qwen3.8-Flash-Next81.6%estimated ± 1.6 pp, high confidence
23GPT-6 Luna81.6%estimated ± 1.6 pp, high confidence
24Seed 2.1 Turbo81.4%estimated ± 0.6 pp, high confidence
25Apodex 1.181.3%estimated ± 1.6 pp, high confidence
26Apodex 1.1 Mini81.3%estimated ± 1.6 pp, high confidence
27Gemini 3.5 Flash-Lite81.1%estimated ± 1.6 pp, high confidence
28Ling 3.0 Flash VL81.1%estimated ± 1.6 pp, high confidence
29Claude Opus 4.7 (Adaptive)81.0%estimated ± 1.6 pp, high confidence
30Gemini 3 Flash80.9%estimated ± 1.6 pp, high confidence
31GPT-5.280.8%estimated ± 0.6 pp, high confidence
32GPT-5.3 Codex80.8%estimated ± 1.6 pp, high confidence
33dots3-note Preview80.4%estimated ± 0.6 pp, high confidence
34Qwen3.5 397B80.3%estimated ± 0.6 pp, high confidence
35Qwen3.7 Plus80.3%estimated ± 0.6 pp, high confidence
36Qwen3.6 Plus80.1%estimated ± 0.6 pp, high confidence
37Kimi K2.680.1%measured
38Claude Sonnet 580.1%estimated ± 1.6 pp, high confidence
39DeepSeek V4.1 Flash79.9%estimated ± 1.6 pp, high confidence
40Kimi K2.579.8%estimated ± 0.6 pp, high confidence
41Kimi K2.5 (Reasoning)79.8%estimated ± 0.6 pp, high confidence
42Claude Opus 4.779.5%estimated ± 1.6 pp, high confidence
43Mistral Large 479.5%estimated ± 1.6 pp, high confidence
44GPT-5.6 Luna79.5%measured
45GPT-5.2-Codex79.5%estimated ± 1.6 pp, high confidence
46Qwen3.8-27B79.5%estimated ± 1.6 pp, high confidence
47Grok 4.379.5%estimated ± 0.6 pp, high confidence
48MiniMax M379.5%estimated ± 0.6 pp, high confidence
49Pareto 26.979.4%estimated ± 0.6 pp, high confidence
50MiMo-V2.579.3%estimated ± 0.6 pp, high confidence
51GPT-5.179.0%estimated ± 1.6 pp, high confidence
52Claude Opus 4.6 (Adaptive)78.9%estimated ± 1.6 pp, high confidence
53Step 3.7 Flash78.8%estimated ± 1.6 pp, high confidence
54Claude Opus 4.678.7%estimated ± 0.6 pp, high confidence
55Qwen3.5-122B-A10B78.7%estimated ± 1.6 pp, high confidence
56Qwen3.5-27B78.7%estimated ± 1.6 pp, high confidence
57Gemini 2.5 Pro78.6%estimated ± 1.6 pp, high confidence
58Gemma 4 31B78.3%estimated ± 0.6 pp, high confidence
59GPT-5 (medium)78.2%estimated ± 1.6 pp, high confidence
60GPT-5 (high)78.1%estimated ± 1.6 pp, high confidence
61Claude Opus 4.5 Thinking78.0%estimated ± 1.6 pp, high confidence
62GPT-5.4 mini78.0%measured
63Step 5 Preview77.4%estimated ± 0.6 pp, high confidence
64MiMo-V2.6-Flash77.4%estimated ± 1.6 pp, high confidence
65Qwen3.6-27B77.2%estimated ± 0.6 pp, high confidence
66GLM-5V-Turbo77.2%estimated ± 1.6 pp, high confidence
67Qwen3.5-35B-A3B77.2%estimated ± 1.6 pp, high confidence
68GPT-5.1-Codex77.0%estimated ± 1.6 pp, high confidence
69GPT-5.1-Codex-Max77.0%estimated ± 1.6 pp, high confidence
70Qwen3.6-35B-A3B76.8%estimated ± 0.6 pp, high confidence
71Grok 4.2076.7%estimated ± 0.6 pp, high confidence
72Claude Sonnet 4.675.8%estimated ± 1.6 pp, high confidence
73Inkling-Small75.6%estimated ± 0.6 pp, high confidence
74Muse Glimmer 30B75.6%estimated ± 0.6 pp, high confidence
75o375.5%estimated ± 1.6 pp, high confidence
76Gemma 4 26B A4B75.4%estimated ± 0.6 pp, high confidence
77MiMo-V2-Omni75.3%estimated ± 1.6 pp, high confidence
78Inkling75.1%estimated ± 0.6 pp, high confidence
79Grok 474.6%estimated ± 1.6 pp, high confidence
80Claude 4.1 Opus Thinking74.0%estimated ± 1.6 pp, high confidence
81Interfaze Beta73.1%estimated ± 0.6 pp, high confidence
82Claude Opus 4.572.7%estimated ± 0.6 pp, high confidence
83Gemini 2.5 Flash72.3%estimated ± 1.6 pp, high confidence
84Mistral Medium 3.5 128B71.9%estimated ± 1.6 pp, medium confidence
85Gemma 4 12B71.5%estimated ± 0.6 pp, high confidence
86Grok 4.1 Fast (Reasoning)70.7%estimated ± 1.6 pp, medium confidence
87Claude 4 Sonnet70.1%estimated ± 1.6 pp, medium confidence
88Llama 4 Maverick69.8%estimated ± 1.6 pp, medium confidence
89Grok 4 Fast (Reasoning)69.6%estimated ± 1.6 pp, medium confidence
90GPT-5.4 nano69.5%measured
91GPT-4.169.2%estimated ± 1.6 pp, medium confidence
92Qwen3-Omni-30B-A3B-Thinking68.4%estimated ± 1.6 pp, medium confidence
93Command A+67.5%estimated ± 0.6 pp, medium confidence
94GPT-4.1 mini67.3%estimated ± 1.6 pp, medium confidence
95Mistral Small 465.8%estimated ± 1.6 pp, medium confidence
96Mistral Small 4 (Reasoning)65.8%estimated ± 1.6 pp, medium confidence
97Mistral Large 365.0%estimated ± 1.6 pp, medium confidence
98Qwen3-Omni-30B-A3B-Instruct64.8%estimated ± 1.6 pp, medium confidence
99Gemini 1.5 Pro64.4%estimated ± 1.6 pp, medium confidence
100Nemotron 3 Nano Omni 30B A3B63.0%estimated ± 1.6 pp, medium confidence
101Mistral Medium 362.8%estimated ± 1.6 pp, medium confidence
102Llama 4 Scout62.7%estimated ± 1.6 pp, medium confidence
103Qwen3.5 397B (Reasoning)62.6%estimated ± 1.6 pp, medium confidence
104Gemma 4 E4B61.5%estimated ± 1.6 pp, medium confidence
105LFM2.5-VL-3B60.9%estimated ± 0.6 pp, medium confidence
106Grok 4.1 Fast59.0%estimated ± 1.6 pp, medium confidence
107Gemma 3 27B58.6%estimated ± 1.6 pp, medium confidence
108Gemma 4 E2B55.6%estimated ± 1.6 pp, medium confidence
109Nova Pro55.4%estimated ± 1.6 pp, medium confidence
110GPT-4o mini52.8%estimated ± 1.6 pp, medium confidence
111GPT-4.1 nano51.5%estimated ± 1.6 pp, medium confidence
112Claude 3 Haiku42.0%estimated ± 1.6 pp, medium confidence
113LFM2.5-VL-1.6B-Extract37.3%estimated ± 1.6 pp, medium confidence
114Phi-4 Multimodal Instruct22.3%estimated ± 1.6 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General