benchgap
Vision & documents

MMMU leaderboard

As of 2026-10-07, the highest measured score on MMMU is 86.0% by Qwen3.6 Plus. 108 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.594.0%estimated ± 0.9 pp, medium confidence
2GPT-6 Astra93.3%estimated ± 0.9 pp, medium confidence
3GPT-6.1 Sol92.5%estimated ± 0.9 pp, medium confidence
4Gemini 3.8 Flash92.2%estimated ± 0.9 pp, medium confidence
5Gemini 3.7 Flash92.1%estimated ± 0.9 pp, medium confidence
6Qwen3.8-Omni-Flash91.7%estimated ± 6.6 pp, low confidence
7Claude Opus 591.4%estimated ± 0.9 pp, medium confidence
8Gemini 3.5 Flash91.0%estimated ± 0.9 pp, medium confidence
9GPT-5.6 Sol90.3%estimated ± 0.9 pp, medium confidence
10Gemini 3.6 Flash90.1%estimated ± 0.9 pp, medium confidence
11GPT-6 Sol89.8%estimated ± 0.9 pp, medium confidence
12Qwen3.8 Max Preview89.7%estimated ± 0.9 pp, medium confidence
13Gemini 3.1 Pro89.4%estimated ± 0.9 pp, medium confidence
14GPT-5.6 Terra87.9%estimated ± 0.9 pp, medium confidence
15Kimi K387.7%estimated ± 0.9 pp, medium confidence
16Muse Spark87.7%estimated ± 0.9 pp, medium confidence
17Qwen3.7 Plus87.7%estimated ± 0.9 pp, medium confidence
18Grok 4.587.6%estimated ± 0.9 pp, medium confidence
19Gemini 3 Pro87.4%estimated ± 0.9 pp, medium confidence
20GPT-5.587.2%estimated ± 0.9 pp, medium confidence
21Qwen3.8-Flash-Next87.1%estimated ± 0.9 pp, medium confidence
22GPT-6 Luna87.0%estimated ± 0.9 pp, medium confidence
23Qwen3.8 Max86.9%estimated ± 1.1 pp, low confidence
24Kimi K2.686.7%estimated ± 0.9 pp, medium confidence
25Apodex 1.186.6%estimated ± 0.9 pp, medium confidence
26Apodex 1.1 Mini86.6%estimated ± 0.9 pp, medium confidence
27Seed 2.1 Pro86.5%estimated ± 1.1 pp, low confidence
28Gemini 3.5 Flash-Lite86.4%estimated ± 0.9 pp, medium confidence
29Ling 3.0 Flash VL86.4%estimated ± 0.9 pp, medium confidence
30Claude Opus 4.7 (Adaptive)86.2%estimated ± 0.9 pp, medium confidence
31Gemini 3 Flash86.0%estimated ± 0.9 pp, medium confidence
32GPT-5.6 Luna86.0%estimated ± 0.9 pp, medium confidence
33MiniMax M386.0%estimated ± 0.9 pp, medium confidence
34Qwen3.6 Plus86.0%measured
35GPT-5.3 Codex85.9%estimated ± 0.9 pp, medium confidence
36GPT-5.485.9%estimated ± 0.9 pp, medium confidence
37Seed 2.1 Turbo85.7%estimated ± 1.1 pp, low confidence
38Grok 4.385.6%estimated ± 0.9 pp, medium confidence
39GPT-5.285.3%estimated ± 1.1 pp, low confidence
40dots3-note Preview85.1%estimated ± 1.1 pp, low confidence
41Claude Sonnet 584.9%estimated ± 0.9 pp, high confidence
42DeepSeek V4.1 Flash84.7%estimated ± 0.9 pp, high confidence
43Pareto 26.984.4%estimated ± 1.1 pp, medium confidence
44MiMo-V2.584.4%estimated ± 1.1 pp, medium confidence
45Claude Opus 4.784.2%estimated ± 0.9 pp, high confidence
46Mistral Large 484.2%estimated ± 0.9 pp, high confidence
47Step 5 Preview84.2%estimated ± 0.9 pp, high confidence
48GPT-5.2-Codex84.1%estimated ± 0.9 pp, high confidence
49Qwen3.8-27B84.1%estimated ± 0.9 pp, high confidence
50Qwen3.5-122B-A10B83.9%measured
51GPT-5.183.4%estimated ± 0.9 pp, high confidence
52Claude Opus 4.6 (Adaptive)83.3%estimated ± 0.9 pp, high confidence
53Kimi K2.583.3%estimated ± 0.9 pp, high confidence
54Kimi K2.5 (Reasoning)83.3%estimated ± 0.9 pp, high confidence
55Step 3.7 Flash83.2%estimated ± 0.9 pp, high confidence
56Gemini 2.5 Pro82.9%estimated ± 0.9 pp, high confidence
57Qwen3.6-27B82.9%measured
58Grok 4.2082.7%estimated ± 1.1 pp, medium confidence
59GPT-5 (medium)82.4%estimated ± 0.9 pp, high confidence
60Muse Glimmer 30B82.4%estimated ± 0.9 pp, high confidence
61GPT-5 (high)82.3%estimated ± 0.9 pp, high confidence
62Qwen3.5-27B82.3%measured
63Claude Opus 4.5 Thinking82.2%estimated ± 0.9 pp, high confidence
64Inkling-Small82.2%estimated ± 0.9 pp, high confidence
65Inkling81.8%estimated ± 0.9 pp, high confidence
66Gemma 4 31B81.7%estimated ± 0.9 pp, high confidence
67Qwen3.6-35B-A3B81.7%measured
68GPT-5.4 mini81.6%estimated ± 0.9 pp, high confidence
69MiMo-V2.6-Flash81.5%estimated ± 0.9 pp, high confidence
70Qwen3.5-35B-A3B81.4%measured
71GLM-5V-Turbo81.2%estimated ± 0.9 pp, high confidence
72Claude Opus 4.681.0%estimated ± 0.9 pp, high confidence
73GPT-5.1-Codex81.0%estimated ± 0.9 pp, high confidence
74GPT-5.1-Codex-Max81.0%estimated ± 0.9 pp, high confidence
75Claude Opus 4.580.0%estimated ± 0.9 pp, high confidence
76Claude Sonnet 4.679.6%estimated ± 0.9 pp, high confidence
77o379.2%estimated ± 0.9 pp, high confidence
78MiMo-V2-Omni79.1%estimated ± 0.9 pp, high confidence
79Gemma 4 12B78.9%estimated ± 0.9 pp, high confidence
80Gemma 4 26B A4B78.6%estimated ± 0.9 pp, high confidence
81Grok 478.3%estimated ± 0.9 pp, high confidence
82Claude 4.1 Opus Thinking77.7%estimated ± 0.9 pp, high confidence
83Gemini 2.5 Flash76.2%estimated ± 0.9 pp, high confidence
84GPT-5.4 nano76.1%estimated ± 0.9 pp, high confidence
85Mistral Medium 3.5 128B75.8%estimated ± 0.9 pp, high confidence
86Command A+75.1%measured
87Grok 4.1 Fast (Reasoning)74.9%estimated ± 0.9 pp, high confidence
88Claude 4 Sonnet74.5%estimated ± 0.9 pp, high confidence
89Llama 4 Maverick74.3%estimated ± 0.9 pp, high confidence
90Grok 4 Fast (Reasoning)74.2%estimated ± 0.9 pp, high confidence
91GPT-4.173.9%estimated ± 0.9 pp, high confidence
92Qwen3-Omni-30B-A3B-Thinking73.4%estimated ± 0.9 pp, high confidence
93GPT-4.1 mini72.8%estimated ± 0.9 pp, high confidence
94Mistral Small 472.0%estimated ± 0.9 pp, high confidence
95Mistral Small 4 (Reasoning)72.0%estimated ± 0.9 pp, high confidence
96Mistral Large 371.7%estimated ± 0.9 pp, high confidence
97Qwen3-Omni-30B-A3B-Instruct71.6%estimated ± 0.9 pp, high confidence
98Gemini 1.5 Pro71.4%estimated ± 0.9 pp, high confidence
99Mistral Medium 370.8%estimated ± 0.9 pp, medium confidence
100Llama 4 Scout70.8%estimated ± 0.9 pp, medium confidence
101Nemotron 3 Nano Omni 30B A3B70.8%measured
102Qwen3.5 397B70.8%estimated ± 0.9 pp, medium confidence
103Qwen3.5 397B (Reasoning)70.8%estimated ± 0.9 pp, medium confidence
104Gemma 4 E4B70.4%estimated ± 0.9 pp, medium confidence
105Grok 4.1 Fast69.8%estimated ± 0.9 pp, medium confidence
106Gemma 3 27B69.7%estimated ± 0.9 pp, medium confidence
107Gemma 4 E2B69.2%estimated ± 0.9 pp, medium confidence
108Nova Pro69.2%estimated ± 0.9 pp, medium confidence
109GPT-4o mini68.9%estimated ± 0.9 pp, medium confidence
110Ternary Bonsai 2 27B68.8%estimated ± 6.6 pp, low confidence
111GPT-4.1 nano68.8%estimated ± 0.9 pp, medium confidence
112Claude 3 Haiku68.4%estimated ± 0.9 pp, medium confidence
113LFM2.5-VL-1.6B-Extract68.3%estimated ± 0.9 pp, medium confidence
114Phi-4 Multimodal Instruct68.3%estimated ± 0.9 pp, medium confidence
115LFM2.5-VL-3B48.4%measured
116ZAYA1-VL-8B46.0%measured
117Interfaze Beta46.0%estimated ± 0.4 pp, low confidence
118North Micro Vision Instruct37.5%estimated ± 6.6 pp, low confidence
119LFM2.5-VL-450M32.7%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General