benchgap
Vision & documents

MMMU-Pro leaderboard

As of 2026-10-07, the highest measured score on MMMU-Pro is 83.9% by Gemini 3.1 Pro. 88 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 3.1 Pro83.9%measured
2Gemini 3.5 Flash83.6%measured
3GPT-5.6 Sol83.0%measured
4Claude Opus 5.582.5%estimated ± 4.4 pp, low confidence
5Qwen3.8 Max82.3%measured
6Claude Mythos 582.3%estimated ± 1.4 pp, low confidence
7Gemini 3.8 Flash82.2%estimated ± 1.4 pp, medium confidence
8GPT-6 Astra81.9%estimated ± 2.7 pp, medium confidence
9GPT-6.1 Sol81.8%estimated ± 4.4 pp, low confidence
10Step 3.7 Flash81.8%estimated ± 2.4 pp, high confidence
11Kimi K381.6%measured
12Seed 2.1 Pro81.6%measured
13Qwen3.8-Flash-Next81.5%estimated ± 1.4 pp, medium confidence
14GLM-5.3-Flash81.4%estimated ± 3.2 pp, high confidence
15Gemini 3.7 Flash81.4%estimated ± 1.4 pp, medium confidence
16Claude Opus 581.3%estimated ± 4.4 pp, low confidence
17GPT-5.481.2%measured
18GPT-5.581.2%measured
19Muse Spark 1.181.1%estimated ± 3.2 pp, high confidence
20Gemini 3 Pro81.0%measured
21GPT-5.6 Terra80.7%measured
22Gemini 3.6 Flash80.6%estimated ± 4.4 pp, medium confidence
23GPT-6 Sol80.5%estimated ± 4.4 pp, medium confidence
24Sakana Fugu-Ultra80.5%estimated ± 3.2 pp, high confidence
25Qwen3.8 Max Preview80.4%estimated ± 4.4 pp, medium confidence
26Muse Spark80.4%measured
27Qwen3.8-27B80.3%estimated ± 1.4 pp, medium confidence
28Seed 2.1 Turbo80.1%measured
29Qwen3.8-Omni-Flash79.9%estimated ± 1.4 pp, medium confidence
30Sakana Fugu79.9%estimated ± 3.2 pp, high confidence
31GPT-5.279.5%measured
32Kimi K2.679.4%measured
33Grok 4.579.4%estimated ± 4.4 pp, medium confidence
34dots3-note Preview79.1%measured
35GPT-6 Luna79.1%estimated ± 4.4 pp, medium confidence
36Qwen3.5 397B79.0%measured
37Qwen3.7 Plus79.0%measured
38Apodex 1.178.8%estimated ± 4.4 pp, medium confidence
39Apodex 1.1 Mini78.8%estimated ± 4.4 pp, medium confidence
40Qwen3.6 Plus78.8%measured
41Gemini 3.5 Flash-Lite78.8%estimated ± 4.4 pp, medium confidence
42Ling 3.0 Flash VL78.8%estimated ± 4.4 pp, medium confidence
43Gemini 3 Flash78.6%estimated ± 4.4 pp, medium confidence
44GPT-5.3 Codex78.5%estimated ± 4.4 pp, medium confidence
45Kimi K2.578.5%measured
46Kimi K2.5 (Reasoning)78.5%measured
47GPT-5.6 Luna78.4%measured
48Holo2-235B-A22B78.3%estimated ± 2.7 pp, high confidence
49Grok 4.378.1%measured
50MiniMax M378.1%measured
51Pareto 26.978.0%measured
52Qwen3.5-122B-A10B77.9%estimated ± 2.2 pp, high confidence
53MiMo-V2.577.9%measured
54DeepSeek V4.1 Flash77.9%estimated ± 4.4 pp, medium confidence
55Qwen3.5-27B77.8%estimated ± 2.2 pp, high confidence
56Claude Opus 4.777.6%estimated ± 4.4 pp, medium confidence
57Mistral Large 477.6%estimated ± 4.4 pp, medium confidence
58GPT-5.2-Codex77.5%estimated ± 4.4 pp, medium confidence
59Holo2-30B-A3B77.3%estimated ± 2.7 pp, high confidence
60Claude Opus 4.677.3%measured
61GPT-5.177.2%estimated ± 4.4 pp, medium confidence
62Claude Opus 4.6 (Adaptive)77.1%estimated ± 4.4 pp, medium confidence
63Claude Sonnet 4.677.1%estimated ± 3.2 pp, high confidence
64Gemini 2.5 Pro76.9%estimated ± 4.4 pp, medium confidence
65Gemma 4 31B76.9%measured
66Claude Opus 4.7 (Adaptive)76.7%estimated ± 1.4 pp, medium confidence
67GPT-5 (medium)76.6%estimated ± 4.4 pp, medium confidence
68Qwen3.5-35B-A3B76.6%estimated ± 2.2 pp, high confidence
69GPT-5.4 mini76.6%measured
70GPT-5 (high)76.6%estimated ± 4.4 pp, medium confidence
71Claude Opus 4.5 Thinking76.5%estimated ± 4.4 pp, medium confidence
72MiMo-V2.6-Flash76.1%estimated ± 4.4 pp, medium confidence
73Step 5 Preview76.0%measured
74GLM-5V-Turbo75.9%estimated ± 4.4 pp, medium confidence
75Qwen3.6-27B75.8%measured
76GPT-5.1-Codex75.8%estimated ± 4.4 pp, medium confidence
77GPT-5.1-Codex-Max75.8%estimated ± 4.4 pp, medium confidence
78Holo2-8B75.5%estimated ± 2.7 pp, high confidence
79Gemini 3.1 Flash-Lite75.4%estimated ± 3.2 pp, high confidence
80Qwen3.6-35B-A3B75.3%measured
81Grok 4.2075.2%measured
82Holo2-4B75.0%estimated ± 2.7 pp, high confidence
83o374.7%estimated ± 4.4 pp, medium confidence
84MiMo-V2-Omni74.6%estimated ± 4.4 pp, medium confidence
85Claude Opus 4.874.4%estimated ± 1.4 pp, medium confidence
86Grok 474.1%estimated ± 4.4 pp, medium confidence
87Inkling-Small74.0%measured
88Muse Glimmer 30B74.0%measured
89Gemma 4 26B A4B73.8%measured
90Claude Sonnet 573.7%estimated ± 1.4 pp, low confidence
91Claude 4.1 Opus Thinking73.6%estimated ± 4.4 pp, medium confidence
92Inkling73.5%measured
93Gemini 2.5 Flash72.5%estimated ± 4.4 pp, medium confidence
94Mistral Medium 3.5 128B72.2%estimated ± 4.4 pp, medium confidence
95Grok 4.1 Fast (Reasoning)71.4%estimated ± 4.4 pp, medium confidence
96Interfaze Beta71.1%measured
97Claude 4 Sonnet71.0%estimated ± 4.4 pp, medium confidence
98Llama 4 Maverick70.8%estimated ± 4.4 pp, medium confidence
99Grok 4 Fast (Reasoning)70.7%estimated ± 4.4 pp, medium confidence
100Claude Opus 4.570.6%measured
101GPT-4.170.4%estimated ± 4.4 pp, medium confidence
102Qwen3-Omni-30B-A3B-Thinking69.9%estimated ± 4.4 pp, medium confidence
103GPT-4.1 mini69.1%estimated ± 4.4 pp, medium confidence
104Gemma 4 12B69.1%measured
105Mistral Small 468.2%estimated ± 4.4 pp, medium confidence
106Mistral Small 4 (Reasoning)68.2%estimated ± 4.4 pp, medium confidence
107Mistral Large 367.6%estimated ± 4.4 pp, medium confidence
108Qwen3-Omni-30B-A3B-Instruct67.5%estimated ± 4.4 pp, medium confidence
109Gemini 1.5 Pro67.2%estimated ± 4.4 pp, medium confidence
110Mistral Medium 366.2%estimated ± 4.4 pp, medium confidence
111Llama 4 Scout66.1%estimated ± 4.4 pp, medium confidence
112GPT-5.4 nano66.1%measured
113Qwen3.5 397B (Reasoning)66.0%estimated ± 4.4 pp, medium confidence
114Gemma 4 E4B65.4%estimated ± 4.4 pp, low confidence
115Grok 4.1 Fast63.8%estimated ± 4.4 pp, low confidence
116Gemma 3 27B63.5%estimated ± 4.4 pp, low confidence
117Command A+63.0%measured
118Ternary Bonsai 2 27B61.8%estimated ± 11.4 pp, low confidence
119Gemma 4 E2B61.7%estimated ± 4.4 pp, low confidence
120Nova Pro61.5%estimated ± 4.4 pp, low confidence
121GPT-4o mini60.0%estimated ± 4.4 pp, low confidence
122GPT-4.1 nano59.2%estimated ± 4.4 pp, low confidence
123Nemotron 3 Nano Omni 30B A3B57.2%estimated ± 2.4 pp, medium confidence
124Claude 3 Haiku53.7%estimated ± 4.4 pp, low confidence
125LFM2.5-VL-1.6B-Extract51.1%estimated ± 4.4 pp, low confidence
126Phi-4 Multimodal Instruct43.3%estimated ± 4.4 pp, low confidence
127LFM2.5-VL-3B30.5%measured
128ZAYA1-VL-8B28.2%estimated ± 2.4 pp, low confidence
129North Micro Vision Instruct22.6%estimated ± 11.4 pp, low confidence
130LFM2.5-VL-450M16.0%estimated ± 2.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General