benchgap
Knowledge & reasoning

MMLU-Pro leaderboard

As of 2026-10-07, the highest measured score on MMLU-Pro is 89.6% by Qwen3.7 Max. 190 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.7 Max89.6%measured
2Claude Opus 4.589.5%measured
3Qwen3.6 Plus88.5%measured
4Qwen3.7 Plus88.5%measured
5Pareto 26.10 Preview88.5%estimated ± 5.4 pp, medium confidence
6Gemini 3 Pro Deep Think88.5%estimated ± 5.8 pp, low confidence
7GLM-5.3-Flash88.2%estimated ± 3.9 pp, high confidence
8Qwen3.5 397B87.8%measured
9GPT-6 Astra87.6%estimated ± 3.0 pp, medium confidence
10DeepSeek V4 Pro 081387.5%measured
11GPT-5.6 Sol87.4%estimated ± 3.0 pp, medium confidence
12Kimi K2.587.1%measured
13Kimi K2.5 (Reasoning)87.1%measured
14GPT-5.6 Terra87.0%estimated ± 3.0 pp, medium confidence
15GPT-5.286.9%estimated ± 3.0 pp, high confidence
16GPT-5.6 Luna86.9%estimated ± 3.0 pp, high confidence
17Gemini 3.5 Flash86.8%estimated ± 3.0 pp, high confidence
18Nemotron 3 Ultra86.8%measured
19Qwen3.5-122B-A10B86.7%measured
20Qwen3.8-Omni-Flash86.6%estimated ± 3.0 pp, high confidence
21DeepSeek V4.1 Flash86.5%estimated ± 3.0 pp, high confidence
22Kimi K2.686.5%estimated ± 3.0 pp, high confidence
23Grok 4.386.4%estimated ± 3.0 pp, high confidence
24dots3-note Preview86.3%estimated ± 3.6 pp, high confidence
25Agents-A186.3%estimated ± 3.6 pp, high confidence
26Interfaze Beta86.3%estimated ± 3.0 pp, high confidence
27Solar Pro 486.3%measured
28DeepSeek V4 Flash 073186.2%measured
29Qwen3.6-27B86.2%measured
30Solar Open 286.2%measured
31Qwen3.5-27B86.1%measured
32Claude Opus 5.586.1%estimated ± 2.5 pp, low confidence
33Claude Fable 5.186.0%estimated ± 2.5 pp, low confidence
34Claude Mythos 586.0%estimated ± 2.5 pp, low confidence
35Claude Sonnet 5.586.0%estimated ± 2.5 pp, low confidence
36Claude Opus 586.0%estimated ± 2.5 pp, low confidence
37Muse Spark 1.185.9%estimated ± 2.5 pp, low confidence
38Sakana Fugu-Ultra85.8%estimated ± 2.5 pp, low confidence
39Claude Opus 4.885.8%estimated ± 2.5 pp, low confidence
40Pareto 26.985.8%estimated ± 2.5 pp, low confidence
41Sakana Fugu85.8%estimated ± 2.5 pp, low confidence
42Claude Opus 4.7 (Adaptive)85.8%estimated ± 2.5 pp, low confidence
43Claude Haiku 5.585.8%estimated ± 2.5 pp, low confidence
44Claude Fable 585.7%estimated ± 3.5 pp, medium confidence
45GPT-6.1 Sol85.7%estimated ± 3.5 pp, medium confidence
46Gemini 3.7 Flash85.7%estimated ± 3.5 pp, medium confidence
47Gemini 3.8 Flash85.7%estimated ± 3.5 pp, medium confidence
48Gemini 3 Pro85.7%estimated ± 3.5 pp, medium confidence
49GPT-5.3 Codex85.7%estimated ± 3.5 pp, medium confidence
50GPT-6 Sol85.7%estimated ± 3.5 pp, medium confidence
51Grok 4.585.7%estimated ± 3.5 pp, medium confidence
52Gemini 3.6 Flash85.7%estimated ± 3.5 pp, medium confidence
53Gemini 4 Argon85.7%estimated ± 3.5 pp, medium confidence
54Grok 4.685.7%estimated ± 3.5 pp, high confidence
55Claude Opus 4.6 (Adaptive)85.7%estimated ± 3.5 pp, high confidence
56Grok 4.785.7%estimated ± 3.5 pp, high confidence
57Claude Opus 4.5 Thinking85.7%estimated ± 3.5 pp, high confidence
58Gemini 3 Flash85.7%estimated ± 3.5 pp, high confidence
59Muse Spark 1.285.7%estimated ± 3.5 pp, high confidence
60Claude Opus 4.785.7%estimated ± 3.5 pp, high confidence
61GPT-6 Luna85.7%estimated ± 3.5 pp, high confidence
62Muse Spark 1.385.7%estimated ± 3.5 pp, high confidence
63Step 5 Preview85.7%estimated ± 3.5 pp, high confidence
64GPT-5.2-Codex85.7%estimated ± 3.5 pp, high confidence
65Grok 485.7%estimated ± 3.5 pp, high confidence
66GPT-5 (high)85.7%estimated ± 3.5 pp, high confidence
67GPT-5.1-Codex85.7%estimated ± 3.5 pp, high confidence
68GPT-5.1-Codex-Max85.7%estimated ± 3.5 pp, high confidence
69Kimi K2.7 Code85.7%estimated ± 3.5 pp, high confidence
70GPT-5 (medium)85.7%estimated ± 3.5 pp, high confidence
71Gemini 3.1 Pro85.7%estimated ± 2.5 pp, low confidence
72o385.7%estimated ± 3.5 pp, high confidence
73Qwen 3.6 Max (preview)85.7%estimated ± 3.5 pp, high confidence
74GPT-5.185.7%estimated ± 3.5 pp, high confidence
75MiMo-V2.6-Pro85.7%estimated ± 3.5 pp, high confidence
76GLM-5.385.7%estimated ± 3.5 pp, high confidence
77Ornith-1.5-397B85.7%estimated ± 2.5 pp, low confidence
78Hy385.7%estimated ± 3.5 pp, high confidence
79Apodex 1.185.7%estimated ± 3.5 pp, high confidence
80Apodex 1.1 Mini85.7%estimated ± 3.5 pp, high confidence
81Qwen3.8 Max Preview85.7%estimated ± 3.5 pp, high confidence
82DeepSeek-R185.7%estimated ± 3.5 pp, high confidence
83GLM-585.7%measured
84Qwen3.8 Max85.7%estimated ± 2.5 pp, low confidence
85Kimi K385.7%estimated ± 2.5 pp, low confidence
86Gemini 3.5 Flash-Lite85.7%estimated ± 3.5 pp, high confidence
87Hy4 preview85.7%estimated ± 2.5 pp, low confidence
88GLM-5V-Turbo85.7%estimated ± 3.5 pp, high confidence
89Claude Sonnet 585.7%estimated ± 2.5 pp, low confidence
90Ling 3.1 Flash85.7%estimated ± 3.5 pp, high confidence
91DeepSeek V3.1 (Reasoning)85.7%estimated ± 3.5 pp, high confidence
92GPT-5.5 Pro85.7%estimated ± 2.5 pp, low confidence
93Muse Spark85.7%estimated ± 2.5 pp, low confidence
94GPT-5.4 Pro85.7%estimated ± 2.5 pp, low confidence
95GLM-5-Turbo85.7%estimated ± 3.5 pp, high confidence
96Kimi K285.6%estimated ± 3.5 pp, high confidence
97MiMo-V2.6-Flash85.6%estimated ± 3.5 pp, high confidence
98Muse Glimmer 30B85.6%estimated ± 3.5 pp, high confidence
99GPT-5.585.6%estimated ± 2.5 pp, low confidence
100MiniMax M2.785.6%estimated ± 3.5 pp, high confidence
101MiMo-V2-Pro85.6%estimated ± 3.5 pp, high confidence
102Hy3 Preview85.6%estimated ± 3.0 pp, high confidence
103GLM-5.285.6%estimated ± 2.5 pp, low confidence
104Gemini 2.5 Flash85.6%estimated ± 3.5 pp, high confidence
105Mistral Large 485.6%estimated ± 3.5 pp, high confidence
106Step 3.7 Flash85.6%estimated ± 3.5 pp, high confidence
107GPT-5.485.6%estimated ± 2.5 pp, medium confidence
108Grok 4.1 Fast (Reasoning)85.6%estimated ± 3.5 pp, high confidence
109Seed 2.1 Pro85.6%estimated ± 6.8 pp, medium confidence
110Mistral Large 385.6%estimated ± 3.5 pp, high confidence
111Llama 4 Maverick85.5%estimated ± 3.5 pp, high confidence
112Mistral Medium 3.5 128B85.5%estimated ± 3.5 pp, high confidence
113Qwen3.5 397B (Reasoning)85.5%estimated ± 3.5 pp, high confidence
114Qwen3 Max85.5%estimated ± 3.5 pp, high confidence
115DeepSeek V3 032485.5%estimated ± 3.5 pp, high confidence
116Nemotron 3 Super 100B85.5%estimated ± 3.5 pp, high confidence
117DeepSeek V3.285.5%estimated ± 3.5 pp, high confidence
118GLM-5.185.5%estimated ± 3.5 pp, high confidence
119Beam85.5%estimated ± 2.5 pp, medium confidence
120Qwen3.8-Flash-Next85.5%estimated ± 2.5 pp, medium confidence
121Grok Code Fast 185.5%estimated ± 3.5 pp, high confidence
122Llama 3.1 405B85.4%estimated ± 3.5 pp, high confidence
123DeepSeek V3.185.4%estimated ± 3.5 pp, high confidence
124Grok 4 Fast (Reasoning)85.4%estimated ± 3.5 pp, high confidence
125Claude 4 Sonnet85.4%estimated ± 3.5 pp, high confidence
126MiMo-V2.5-Pro85.4%estimated ± 2.5 pp, medium confidence
127Trinity-Large-Preview85.4%estimated ± 3.5 pp, high confidence
128Trinity-Large-Thinking85.4%estimated ± 3.5 pp, high confidence
129Qwen3.5-35B-A3B85.3%measured
130Mercury 2.585.3%estimated ± 3.5 pp, high confidence
131GPT-OSS 120B85.3%estimated ± 3.5 pp, high confidence
132Grok 4.2085.3%estimated ± 2.5 pp, medium confidence
133Inkling-Small85.3%estimated ± 2.5 pp, medium confidence
134Mistral Small 485.3%estimated ± 3.5 pp, high confidence
135Mistral Small 4 (Reasoning)85.3%estimated ± 3.5 pp, high confidence
136o3-pro85.2%estimated ± 3.9 pp, high confidence
137Qwen3.8-27B85.2%estimated ± 2.5 pp, medium confidence
138GLM-4.685.2%estimated ± 3.5 pp, high confidence
139Gemma 4 31B85.2%measured
140Qwen3.6-35B-A3B85.2%measured
141Inkling85.2%estimated ± 2.5 pp, medium confidence
142GPT-5.4 mini85.1%estimated ± 2.5 pp, medium confidence
143MAI-Thinking-185.0%measured
144Nemotron Ultra 253B85.0%estimated ± 3.5 pp, high confidence
145Ling 3.0 Flash85.0%estimated ± 3.0 pp, high confidence
146GPT-4o84.9%estimated ± 3.5 pp, high confidence
147Mistral Large 284.9%estimated ± 3.5 pp, high confidence
148Ornith-1.5-35B-A3B84.9%estimated ± 2.5 pp, medium confidence
149MiMo-V2-Flash84.9%measured
150GPT-5.4 nano84.8%estimated ± 2.5 pp, medium confidence
151MiMo-V2-Omni84.8%estimated ± 3.5 pp, high confidence
152Claude 4.1 Opus84.8%estimated ± 4.5 pp, high confidence
153Ultravox v0.6 Llama 3.3 70B84.7%estimated ± 3.5 pp, high confidence
154Ling 3.0 Flash FP884.7%estimated ± 3.0 pp, high confidence
155North Mini Code84.7%estimated ± 3.5 pp, high confidence
156A.X K284.6%estimated ± 3.5 pp, high confidence
157Solar Pro 384.6%estimated ± 3.5 pp, high confidence
158Claude Sonnet 4.584.5%estimated ± 3.0 pp, high confidence
159Mistral Medium 384.5%estimated ± 3.5 pp, high confidence
160Seed 2.1 Turbo84.5%estimated ± 6.8 pp, medium confidence
161Ornith-1.5-9B84.4%estimated ± 2.5 pp, medium confidence
162Gemini 2.5 Pro84.4%estimated ± 3.0 pp, high confidence
163GLM-4.784.3%measured
164Claude 3 Haiku84.2%estimated ± 3.5 pp, high confidence
165Sarvam 105B84.2%estimated ± 3.5 pp, high confidence
166Nemotron 3 Nano 30B84.1%estimated ± 3.5 pp, high confidence
167Grok 4.1 Fast84.0%estimated ± 3.5 pp, high confidence
168Nova Pro83.9%estimated ± 3.5 pp, high confidence
169MiniMax M383.8%estimated ± 3.5 pp, high confidence
170Claude 4.1 Opus Thinking83.7%estimated ± 3.9 pp, high confidence
171K-Exaone83.6%estimated ± 3.5 pp, high confidence
172GLM-4.5-Air83.6%estimated ± 3.5 pp, high confidence
173K-EXAONE 2.083.5%measured
174Solar Pro 283.4%estimated ± 3.5 pp, high confidence
175GPT-OSS 20B83.4%estimated ± 3.5 pp, high confidence
176Quasar 438B83.0%estimated ± 3.5 pp, high confidence
177Qwen3 235B 250783.0%measured
178o1-pro83.0%estimated ± 3.0 pp, high confidence
179Llama 4 Scout82.8%estimated ± 3.5 pp, high confidence
180Gemma 4 26B A4B82.6%measured
181o3-mini82.3%estimated ± 3.0 pp, high confidence
182Qwen3-Omni-30B-A3B-Thinking82.3%estimated ± 3.5 pp, high confidence
183Ling 3.0 Flash VL82.1%estimated ± 3.5 pp, high confidence
184Claude Opus 4.682.0%measured
185Qwen3-Omni-30B-A3B-Instruct82.0%estimated ± 3.5 pp, high confidence
186Exaone 4.0 32B81.8%measured
187Phi-481.8%estimated ± 3.5 pp, high confidence
188o1-preview81.7%estimated ± 3.9 pp, high confidence
189o181.7%estimated ± 3.0 pp, high confidence
190Nemotron 3.5 Lightning 30B A3B NVFP481.6%measured
191Gemma 3 27B80.4%estimated ± 3.5 pp, high confidence
192Sarvam 30B79.8%estimated ± 3.5 pp, high confidence
193LongCat-Flash-Lite-Sparse79.2%measured
194Claude Sonnet 4.679.2%measured
195Ternary Bonsai 2 27B77.9%estimated ± 1.4 pp, high confidence
196Granite 4.2 30B77.6%measured
197Nemotron 3 Nano Omni 30B A3B77.3%measured
198Gemma 4 12B77.2%measured
199GPT-4.176.8%estimated ± 3.0 pp, high confidence
200Celeris-175.9%measured
201DeepSeek V375.9%measured
202GPT-4.1 mini75.5%estimated ± 3.0 pp, high confidence
203DeepSeek R1 Distill Qwen 32B75.2%estimated ± 3.9 pp, high confidence
204ZAYA1-8B74.2%measured
205Granite 4.2 8B74.0%measured
206Gemini 1.5 Pro74.0%estimated ± 3.9 pp, high confidence
207Mellum2-12B-A2.5B-Thinking73.0%estimated ± 1.4 pp, high confidence
208LFM2.5-8B-A1B72.4%estimated ± 3.5 pp, high confidence
209Claude 3.5 Sonnet71.8%estimated ± 3.0 pp, high confidence
210GPT-4 Turbo71.8%estimated ± 4.5 pp, high confidence
211Ling 2.6 Flash71.5%estimated ± 3.0 pp, high confidence
212MiniCPM5-2B70.8%measured
213Command A+70.8%estimated ± 3.5 pp, high confidence
214Claude 3 Opus69.6%estimated ± 3.9 pp, high confidence
215Gemma 4 E4B69.4%measured
216Ling 3.0 Tiny69.3%estimated ± 3.5 pp, high confidence
217ZAYA1-74B-Preview68.1%measured
218Granite 4.2 3B67.8%measured
219GPT-4o mini66.9%estimated ± 3.9 pp, medium confidence
220Qwen2.5 Coder 32B Instruct66.5%estimated ± 3.9 pp, medium confidence
221GPT-4.1 nano62.7%estimated ± 3.0 pp, high confidence
222Phi-4 Multimodal Instruct62.0%estimated ± 3.9 pp, medium confidence
223Mellum2-12B-A2.5B-Instruct60.4%estimated ± 1.4 pp, high confidence
224Gemini 1.0 Pro60.3%estimated ± 3.9 pp, medium confidence
225Gemma 4 E2B60.0%measured
226LFM2.5-VL-1.6B-Extract56.9%estimated ± 3.5 pp, medium confidence
227LLaDA2.2-mini55.2%estimated ± 5.4 pp, medium confidence
228Granite-4.0-H-1B53.5%estimated ± 3.5 pp, medium confidence
229Exaone 4.0 1.2B52.4%estimated ± 3.5 pp, medium confidence
230Soofi S 30B-A3B51.4%measured
231LFM2.5-2.6B48.9%estimated ± 3.5 pp, medium confidence
232MiniCPM5-1B48.9%measured
233Granite-4.0-350M45.9%estimated ± 3.5 pp, medium confidence
234Granite-4.0-H-350M45.3%estimated ± 3.5 pp, medium confidence
235LFM2.5-230M20.3%measured
236LFM2.5-VL-450M19.3%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General