benchgap
Knowledge & reasoning

MMLU leaderboard

As of 2026-10-07, the highest measured score on MMLU is 91.8% by o1. 215 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Haiku 5.5100.0%estimated ± 1.9 pp, low confidence
2GLM-5.3-Flash100.0%estimated ± 1.9 pp, low confidence
3Claude Fable 5.196.6%estimated ± 1.1 pp, low confidence
4Claude Opus 5.596.5%estimated ± 1.1 pp, low confidence
5Claude Fable 596.5%estimated ± 1.1 pp, low confidence
6Claude 4.1 Opus Thinking96.4%estimated ± 1.9 pp, low confidence
7GPT-6 Astra96.3%estimated ± 1.1 pp, low confidence
8GPT-6.1 Sol96.2%estimated ± 1.1 pp, low confidence
9Claude Opus 596.1%estimated ± 1.1 pp, low confidence
10o3-pro96.0%estimated ± 1.9 pp, low confidence
11GPT-5.6 Sol96.0%estimated ± 1.1 pp, low confidence
12GPT-5.595.8%estimated ± 1.1 pp, low confidence
13Gemini 3 Pro95.6%estimated ± 1.1 pp, low confidence
14Gemini 3.7 Flash95.6%estimated ± 1.1 pp, low confidence
15Gemini 3.1 Pro95.5%estimated ± 1.1 pp, low confidence
16Gemini 3.8 Flash95.5%estimated ± 1.1 pp, low confidence
17GPT-6 Sol95.5%estimated ± 1.1 pp, low confidence
18Claude Sonnet 5.595.4%estimated ± 1.1 pp, low confidence
19GPT-5.3 Codex95.3%estimated ± 1.1 pp, low confidence
20Muse Spark 1.195.2%estimated ± 1.1 pp, low confidence
21Gemini 3.5 Flash95.2%estimated ± 1.1 pp, low confidence
22Grok 4.595.1%estimated ± 1.1 pp, low confidence
23GPT-5.495.0%estimated ± 1.1 pp, low confidence
24Gemini 3.6 Flash94.9%estimated ± 1.1 pp, low confidence
25Gemini 4 Argon94.9%estimated ± 1.1 pp, low confidence
26Muse Spark94.9%estimated ± 1.1 pp, low confidence
27DeepSeek V4 Pro 081394.8%estimated ± 1.1 pp, low confidence
28Claude Opus 4.7 (Adaptive)94.8%estimated ± 1.1 pp, low confidence
29Claude Opus 4.894.8%estimated ± 1.1 pp, low confidence
30Grok 4.694.7%estimated ± 1.1 pp, low confidence
31Kimi K394.6%estimated ± 1.1 pp, low confidence
32Grok 4.794.6%estimated ± 1.1 pp, low confidence
33Claude Opus 4.6 (Adaptive)94.5%estimated ± 1.1 pp, low confidence
34GPT-5.6 Terra94.5%estimated ± 1.1 pp, low confidence
35Claude Opus 4.5 Thinking94.5%estimated ± 1.1 pp, low confidence
36DeepSeek V4.1 Flash94.4%estimated ± 1.1 pp, low confidence
37Claude Opus 4.694.4%estimated ± 1.1 pp, low confidence
38Gemini 3 Flash94.4%estimated ± 1.1 pp, low confidence
39Muse Spark 1.294.3%estimated ± 1.1 pp, low confidence
40Claude 4.1 Opus94.2%estimated ± 1.9 pp, low confidence
41Claude Opus 4.794.2%estimated ± 1.1 pp, low confidence
42GPT-5.294.1%estimated ± 1.1 pp, low confidence
43GPT-6 Luna94.0%estimated ± 1.1 pp, low confidence
44Muse Spark 1.394.0%estimated ± 1.1 pp, low confidence
45GPT-5.6 Luna93.9%estimated ± 1.1 pp, low confidence
46Inkling93.7%estimated ± 1.1 pp, low confidence
47Step 5 Preview93.6%estimated ± 1.1 pp, low confidence
48GPT-5.2-Codex93.6%estimated ± 1.1 pp, low confidence
49Claude Opus 4.593.5%estimated ± 1.1 pp, low confidence
50Grok 493.5%estimated ± 1.1 pp, low confidence
51DeepSeek V4 Flash 073193.4%estimated ± 1.1 pp, low confidence
52GPT-5 (high)93.4%estimated ± 1.1 pp, low confidence
53Claude Sonnet 593.4%estimated ± 1.1 pp, low confidence
54GPT-5.1-Codex93.3%estimated ± 1.1 pp, low confidence
55GPT-5.1-Codex-Max93.3%estimated ± 1.1 pp, low confidence
56Kimi K2.7 Code93.3%estimated ± 1.1 pp, low confidence
57GPT-5 (medium)93.3%estimated ± 1.1 pp, low confidence
58Gemini 2.5 Pro93.2%estimated ± 1.1 pp, low confidence
59Claude Sonnet 4.693.1%estimated ± 1.1 pp, low confidence
60o393.1%estimated ± 1.1 pp, low confidence
61Qwen 3.6 Max (preview)92.9%estimated ± 1.1 pp, low confidence
62GPT-5.192.9%estimated ± 1.1 pp, low confidence
63GPT-5.4 mini92.8%estimated ± 1.1 pp, low confidence
64Kimi K2.592.3%estimated ± 1.1 pp, low confidence
65Kimi K2.5 (Reasoning)92.3%estimated ± 1.1 pp, low confidence
66MiMo-V2.6-Pro92.2%estimated ± 1.1 pp, low confidence
67Grok 4.392.2%estimated ± 1.1 pp, low confidence
68GLM-5.392.0%estimated ± 1.1 pp, medium confidence
69o191.8%measured
70Inkling-Small91.8%estimated ± 1.1 pp, medium confidence
71Kimi K2.691.6%estimated ± 1.1 pp, medium confidence
72Hy391.5%estimated ± 1.1 pp, medium confidence
73Apodex 1.191.4%estimated ± 1.1 pp, medium confidence
74Apodex 1.1 Mini91.4%estimated ± 1.1 pp, medium confidence
75Qwen3.8 Max Preview91.4%estimated ± 1.1 pp, medium confidence
76Hy3 Preview91.3%estimated ± 1.1 pp, medium confidence
77Qwen3.7 Max91.2%estimated ± 1.1 pp, medium confidence
78DeepSeek-R191.0%estimated ± 1.1 pp, medium confidence
79Gemini 3.5 Flash-Lite90.7%estimated ± 1.1 pp, medium confidence
80GLM-4.790.6%estimated ± 1.1 pp, medium confidence
81GLM-5V-Turbo90.6%estimated ± 1.1 pp, medium confidence
82Ling 3.1 Flash90.5%estimated ± 1.1 pp, medium confidence
83DeepSeek V3.1 (Reasoning)90.5%estimated ± 1.1 pp, medium confidence
84Sakana Fugu90.5%estimated ± 3.6 pp, low confidence
85Sakana Fugu-Ultra90.5%estimated ± 3.6 pp, low confidence
86Claude Mythos 590.5%estimated ± 3.6 pp, low confidence
87Ornith-1.5-397B90.5%estimated ± 3.6 pp, low confidence
88Qwen3.8 Max90.5%estimated ± 3.6 pp, low confidence
89Hy4 preview90.4%estimated ± 3.6 pp, low confidence
90Qwen3.8-Omni-Flash90.4%estimated ± 3.6 pp, low confidence
91Interfaze Beta90.4%estimated ± 3.6 pp, low confidence
92Ornith-1.5-35B-A3B90.4%estimated ± 3.6 pp, low confidence
93Ornith-1.5-9B90.3%estimated ± 3.6 pp, low confidence
94GLM-5-Turbo90.3%estimated ± 1.1 pp, medium confidence
95Ternary Bonsai 2 27B90.3%estimated ± 3.6 pp, low confidence
96MAI-Thinking-190.2%estimated ± 3.6 pp, low confidence
97GPT-4.190.2%measured
98Claude Sonnet 4.590.2%estimated ± 3.6 pp, low confidence
99Kimi K289.9%estimated ± 1.1 pp, medium confidence
100Qwen3 235B 250789.9%estimated ± 3.6 pp, low confidence
101MiMo-V2.6-Flash89.8%estimated ± 1.1 pp, medium confidence
102Muse Glimmer 30B89.8%estimated ± 1.1 pp, medium confidence
103MiniMax M2.789.7%estimated ± 1.1 pp, medium confidence
104MiMo-V2-Pro89.6%estimated ± 1.1 pp, medium confidence
105Qwen3.6 Plus89.5%estimated ± 1.1 pp, medium confidence
106GLM-589.5%estimated ± 1.1 pp, medium confidence
107Gemini 2.5 Flash89.4%estimated ± 1.1 pp, medium confidence
108ZAYA1-8B89.3%estimated ± 3.6 pp, medium confidence
109Mistral Large 489.3%estimated ± 1.1 pp, medium confidence
110Step 3.7 Flash89.3%estimated ± 1.1 pp, medium confidence
111GPT-5.4 nano89.2%estimated ± 1.1 pp, medium confidence
112DeepSeek V389.1%estimated ± 1.1 pp, medium confidence
113o1-pro89.0%estimated ± 1.9 pp, medium confidence
114Grok 4.1 Fast (Reasoning)89.0%estimated ± 1.1 pp, medium confidence
115Mistral Large 388.9%estimated ± 1.1 pp, medium confidence
116Llama 4 Maverick88.9%estimated ± 1.1 pp, medium confidence
117Mistral Medium 3.5 128B88.8%estimated ± 1.1 pp, medium confidence
118Qwen3.5 397B88.7%estimated ± 1.1 pp, medium confidence
119Qwen3.5 397B (Reasoning)88.7%estimated ± 1.1 pp, medium confidence
120Qwen3.8-Flash-Next88.7%estimated ± 1.1 pp, medium confidence
121Qwen3.5-122B-A10B88.6%estimated ± 1.1 pp, medium confidence
122Qwen3 Max88.6%estimated ± 1.1 pp, medium confidence
123DeepSeek V3 032488.6%estimated ± 1.1 pp, medium confidence
124GLM-5.288.6%estimated ± 1.1 pp, medium confidence
125Nemotron 3 Super 100B88.6%estimated ± 1.1 pp, medium confidence
126DeepSeek V3.288.5%estimated ± 1.1 pp, medium confidence
127GLM-5.188.3%estimated ± 1.1 pp, medium confidence
128Grok Code Fast 188.2%estimated ± 1.1 pp, medium confidence
129Llama 3.1 405B88.1%estimated ± 1.1 pp, medium confidence
130DeepSeek V3.188.0%estimated ± 1.1 pp, medium confidence
131Grok 4 Fast (Reasoning)87.9%estimated ± 1.1 pp, medium confidence
132Claude 4 Sonnet87.8%estimated ± 1.1 pp, medium confidence
133Qwen3.7 Plus87.7%estimated ± 1.1 pp, medium confidence
134Trinity-Large-Thinking87.7%estimated ± 1.1 pp, medium confidence
135MiMo-V2.5-Pro87.6%estimated ± 1.1 pp, medium confidence
136o1-preview87.6%estimated ± 1.9 pp, medium confidence
137GPT-4.1 mini87.5%measured
138Mercury 2.587.4%estimated ± 1.1 pp, medium confidence
139GPT-OSS 120B87.3%estimated ± 1.1 pp, medium confidence
140Mistral Small 487.2%estimated ± 1.1 pp, medium confidence
141Mistral Small 4 (Reasoning)87.2%estimated ± 1.1 pp, medium confidence
142Trinity-Large-Preview87.2%measured
143Nemotron 3 Ultra87.2%estimated ± 1.1 pp, medium confidence
144GLM-4.687.1%estimated ± 1.1 pp, medium confidence
145o3-mini86.9%measured
146Qwen3.5-27B86.6%estimated ± 1.1 pp, medium confidence
147Claude 3.5 Sonnet86.6%estimated ± 3.6 pp, medium confidence
148Nemotron Ultra 253B86.3%estimated ± 1.1 pp, medium confidence
149Qwen3.5-35B-A3B86.3%estimated ± 1.1 pp, medium confidence
150Gemma 4 31B86.2%estimated ± 1.1 pp, medium confidence
151GPT-4o86.1%estimated ± 1.1 pp, medium confidence
152Mistral Large 286.1%estimated ± 1.1 pp, medium confidence
153Qwen3.6-27B85.9%estimated ± 1.1 pp, medium confidence
154Mellum2-12B-A2.5B-Thinking85.8%estimated ± 3.6 pp, medium confidence
155MiMo-V2-Omni85.7%estimated ± 1.1 pp, medium confidence
156ZAYA1-74B-Preview85.6%estimated ± 3.6 pp, medium confidence
157Gemma 4 26B A4B85.6%estimated ± 1.1 pp, medium confidence
158Ultravox v0.6 Llama 3.3 70B85.5%estimated ± 1.1 pp, medium confidence
159North Mini Code85.4%estimated ± 1.1 pp, medium confidence
160Solar Pro 485.4%estimated ± 1.1 pp, medium confidence
161Qwen3.6-35B-A3B85.4%estimated ± 1.1 pp, medium confidence
162LongCat-Flash-Lite-Sparse85.3%measured
163A.X K285.2%estimated ± 1.1 pp, medium confidence
164Solar Pro 385.1%estimated ± 1.1 pp, medium confidence
165Mistral Medium 385.0%estimated ± 1.1 pp, medium confidence
166Ling 3.0 Flash84.9%estimated ± 1.1 pp, medium confidence
167Ling 3.0 Flash FP884.9%estimated ± 1.1 pp, medium confidence
168Claude 3 Haiku84.4%estimated ± 1.1 pp, medium confidence
169Sarvam 105B84.4%estimated ± 1.1 pp, medium confidence
170Nemotron 3 Nano 30B84.2%estimated ± 1.1 pp, medium confidence
171Grok 4.1 Fast84.1%estimated ± 1.1 pp, medium confidence
172Nova Pro83.8%estimated ± 1.1 pp, medium confidence
173MiniMax M383.7%estimated ± 1.1 pp, medium confidence
174K-Exaone83.4%estimated ± 1.1 pp, medium confidence
175GLM-4.5-Air83.3%estimated ± 1.1 pp, medium confidence
176Solar Pro 283.1%estimated ± 1.1 pp, medium confidence
177Claude 3 Opus83.0%estimated ± 1.9 pp, medium confidence
178GPT-OSS 20B83.0%estimated ± 1.1 pp, medium confidence
179Gemma 4 12B82.6%estimated ± 1.1 pp, medium confidence
180Ling 2.6 Flash82.6%estimated ± 1.1 pp, medium confidence
181MiMo-V2-Flash82.6%estimated ± 1.1 pp, medium confidence
182Qwen3.8-27B82.6%estimated ± 1.1 pp, medium confidence
183Quasar 438B82.5%estimated ± 1.1 pp, medium confidence
184DeepSeek R1 Distill Qwen 32B82.3%estimated ± 1.9 pp, medium confidence
185Llama 4 Scout82.2%estimated ± 1.1 pp, medium confidence
186Nemotron 3 Nano Omni 30B A3B82.2%estimated ± 1.1 pp, medium confidence
187Qwen3-Omni-30B-A3B-Thinking81.6%estimated ± 1.1 pp, medium confidence
188Ling 3.0 Flash VL81.3%estimated ± 1.1 pp, medium confidence
189Nemotron 3.5 Lightning 30B A3B NVFP481.3%estimated ± 1.1 pp, medium confidence
190Qwen3-Omni-30B-A3B-Instruct81.2%estimated ± 1.1 pp, medium confidence
191Gemini 1.5 Pro81.2%estimated ± 1.9 pp, medium confidence
192Phi-481.0%estimated ± 1.1 pp, medium confidence
193GPT-4.1 nano80.1%measured
194K-EXAONE 2.079.7%estimated ± 1.1 pp, low confidence
195Gemma 3 27B79.6%estimated ± 1.1 pp, low confidence
196Sarvam 30B79.1%estimated ± 1.1 pp, low confidence
197GPT-4 Turbo78.8%estimated ± 1.9 pp, low confidence
198Qwen2.5 Coder 32B Instruct77.9%estimated ± 1.9 pp, low confidence
199GPT-4o mini77.6%estimated ± 1.9 pp, low confidence
200Granite 4.2 8B76.9%estimated ± 1.1 pp, low confidence
201Celeris-176.6%estimated ± 1.1 pp, low confidence
202Exaone 4.0 32B75.9%estimated ± 1.1 pp, low confidence
203Granite 4.2 30B74.9%estimated ± 1.1 pp, low confidence
204Phi-4 Multimodal Instruct74.6%estimated ± 1.9 pp, low confidence
205LFM2.5-8B-A1B73.5%estimated ± 1.1 pp, low confidence
206Granite 4.2 3B73.0%estimated ± 1.1 pp, low confidence
207Gemini 1.0 Pro72.7%estimated ± 1.9 pp, low confidence
208Command A+72.3%estimated ± 1.1 pp, low confidence
209Gemma 4 E4B71.6%estimated ± 1.1 pp, low confidence
210Ling 3.0 Tiny71.4%estimated ± 1.1 pp, low confidence
211MiniCPM5-2B71.1%estimated ± 1.1 pp, low confidence
212Soofi S 30B-A3B68.9%estimated ± 3.6 pp, low confidence
213Gemma 4 E2B65.7%estimated ± 1.1 pp, low confidence
214LFM2.5-VL-1.6B-Extract62.6%estimated ± 1.1 pp, low confidence
215Mellum2-12B-A2.5B-Instruct62.5%estimated ± 3.6 pp, low confidence
216Granite-4.0-H-1B60.0%estimated ± 1.1 pp, low confidence
217Exaone 4.0 1.2B59.0%estimated ± 1.1 pp, low confidence
218LFM2.5-2.6B55.8%estimated ± 1.1 pp, low confidence
219Granite-4.0-350M52.8%estimated ± 1.1 pp, low confidence
220Granite-4.0-H-350M52.1%estimated ± 1.1 pp, low confidence
221LFM2.5-VL-450M10.8%estimated ± 3.6 pp, low confidence
222LFM2.5-230M10.2%estimated ± 3.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General