benchgap
Knowledge & reasoning

MMLU-Redux leaderboard

As of 2026-10-07, the highest measured score on MMLU-Redux is 96.6% by Claude Opus 4.5. 223 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.198.1%estimated ± 1.1 pp, low confidence
2Claude Opus 5.598.1%estimated ± 1.1 pp, low confidence
3Claude Fable 598.0%estimated ± 1.1 pp, low confidence
4GPT-6 Astra97.9%estimated ± 1.1 pp, low confidence
5GPT-6.1 Sol97.9%estimated ± 1.1 pp, low confidence
6Claude Opus 597.9%estimated ± 1.1 pp, low confidence
7GPT-5.6 Sol97.8%estimated ± 1.1 pp, low confidence
8GPT-5.597.8%estimated ± 1.1 pp, low confidence
9Gemini 3 Pro97.7%estimated ± 1.1 pp, low confidence
10Gemini 3.7 Flash97.6%estimated ± 1.1 pp, low confidence
11Gemini 3.1 Pro97.6%estimated ± 1.1 pp, low confidence
12Gemini 3.8 Flash97.6%estimated ± 1.1 pp, low confidence
13GPT-6 Sol97.6%estimated ± 1.1 pp, low confidence
14Claude Sonnet 5.597.6%estimated ± 1.1 pp, low confidence
15GPT-5.3 Codex97.5%estimated ± 1.1 pp, low confidence
16Muse Spark 1.197.5%estimated ± 1.1 pp, low confidence
17Gemini 3.5 Flash97.5%estimated ± 1.1 pp, low confidence
18Grok 4.597.5%estimated ± 1.1 pp, low confidence
19GPT-5.497.4%estimated ± 1.1 pp, low confidence
20Gemini 3.6 Flash97.4%estimated ± 1.1 pp, low confidence
21Gemini 4 Argon97.4%estimated ± 1.1 pp, low confidence
22Muse Spark97.3%estimated ± 1.1 pp, low confidence
23Claude Opus 4.7 (Adaptive)97.3%estimated ± 1.1 pp, low confidence
24Claude Opus 4.897.3%estimated ± 1.1 pp, low confidence
25Claude Mythos 597.3%estimated ± 2.1 pp, low confidence
26Grok 4.697.3%estimated ± 1.1 pp, low confidence
27Kimi K397.2%estimated ± 1.1 pp, low confidence
28Grok 4.797.2%estimated ± 1.1 pp, low confidence
29Claude Opus 4.6 (Adaptive)97.2%estimated ± 1.1 pp, low confidence
30GPT-5.6 Terra97.2%estimated ± 1.1 pp, low confidence
31Claude Opus 4.5 Thinking97.2%estimated ± 1.1 pp, low confidence
32DeepSeek V4.1 Flash97.1%estimated ± 1.1 pp, low confidence
33Gemini 3 Flash97.1%estimated ± 1.1 pp, low confidence
34Muse Spark 1.297.1%estimated ± 1.1 pp, low confidence
35GPT-5.4 Pro97.0%estimated ± 2.1 pp, low confidence
36Claude Opus 4.797.0%estimated ± 1.1 pp, low confidence
37GPT-5.297.0%estimated ± 1.1 pp, low confidence
38GPT-5.5 Pro97.0%estimated ± 2.1 pp, low confidence
39GPT-6 Luna97.0%estimated ± 1.1 pp, low confidence
40Muse Spark 1.396.9%estimated ± 1.1 pp, low confidence
41Hy4 preview96.9%estimated ± 2.1 pp, low confidence
42GPT-5.6 Luna96.9%estimated ± 1.1 pp, low confidence
43Inkling96.8%estimated ± 1.1 pp, low confidence
44Step 5 Preview96.8%estimated ± 1.1 pp, low confidence
45dots3-note Preview96.8%estimated ± 2.1 pp, low confidence
46GPT-5.2-Codex96.7%estimated ± 1.1 pp, low confidence
47Grok 496.7%estimated ± 1.1 pp, medium confidence
48GPT-5 (high)96.7%estimated ± 1.1 pp, medium confidence
49Claude Sonnet 596.7%estimated ± 1.1 pp, medium confidence
50GPT-5.1-Codex96.6%estimated ± 1.1 pp, medium confidence
51GPT-5.1-Codex-Max96.6%estimated ± 1.1 pp, medium confidence
52Kimi K2.7 Code96.6%estimated ± 1.1 pp, medium confidence
53GPT-5 (medium)96.6%estimated ± 1.1 pp, medium confidence
54Claude Opus 4.596.6%measured
55Gemini 2.5 Pro96.6%estimated ± 1.1 pp, medium confidence
56o396.5%estimated ± 1.1 pp, medium confidence
57Agents-A196.5%estimated ± 2.1 pp, low confidence
58GPT-5.196.4%estimated ± 1.1 pp, medium confidence
59GPT-5.4 mini96.4%estimated ± 1.1 pp, medium confidence
60Ornith-1.5-397B96.3%estimated ± 2.1 pp, low confidence
61Qwen3.8 Max96.2%estimated ± 2.1 pp, low confidence
62MiMo-V2.6-Pro96.1%estimated ± 1.1 pp, medium confidence
63Grok 4.396.1%estimated ± 1.1 pp, medium confidence
64o196.1%estimated ± 1.1 pp, medium confidence
65GLM-5.396.0%estimated ± 1.1 pp, medium confidence
66Inkling-Small95.9%estimated ± 1.1 pp, medium confidence
67Kimi K2.695.9%estimated ± 1.1 pp, medium confidence
68Hy395.8%estimated ± 1.1 pp, medium confidence
69Apodex 1.195.7%estimated ± 1.1 pp, medium confidence
70Apodex 1.1 Mini95.7%estimated ± 1.1 pp, medium confidence
71Qwen3.8 Max Preview95.7%estimated ± 1.1 pp, medium confidence
72Hy3 Preview95.7%estimated ± 1.1 pp, medium confidence
73Qwen 3.6 Max (preview)95.6%estimated ± 0.9 pp, medium confidence
74DeepSeek-R195.6%estimated ± 1.1 pp, medium confidence
75Qwen3.8-Omni-Flash95.6%estimated ± 2.1 pp, medium confidence
76Gemini 3.5 Flash-Lite95.4%estimated ± 1.1 pp, medium confidence
77GLM-5V-Turbo95.4%estimated ± 1.1 pp, medium confidence
78Ling 3.1 Flash95.3%estimated ± 1.1 pp, medium confidence
79DeepSeek V3.1 (Reasoning)95.3%estimated ± 1.1 pp, medium confidence
80GLM-5-Turbo95.2%estimated ± 1.1 pp, medium confidence
81Sakana Fugu95.2%estimated ± 2.1 pp, medium confidence
82Sakana Fugu-Ultra95.2%estimated ± 2.1 pp, medium confidence
83Claude Haiku 5.595.2%estimated ± 2.1 pp, low confidence
84GLM-5.3-Flash95.2%estimated ± 2.1 pp, low confidence
85GPT-4.195.1%estimated ± 1.1 pp, medium confidence
86Kimi K295.0%estimated ± 1.1 pp, medium confidence
87Qwen3.7 Max95.0%measured
88MiMo-V2.6-Flash95.0%estimated ± 1.1 pp, medium confidence
89Muse Glimmer 30B95.0%estimated ± 1.1 pp, medium confidence
90MiniMax M2.794.9%estimated ± 1.1 pp, medium confidence
91Seed 2.1 Pro94.9%estimated ± 0.9 pp, high confidence
92Qwen3.5 397B94.9%measured
93MiMo-V2-Pro94.9%estimated ± 1.1 pp, medium confidence
94Gemini 2.5 Flash94.8%estimated ± 1.1 pp, medium confidence
95Claude 4.1 Opus Thinking94.7%estimated ± 2.1 pp, medium confidence
96Mistral Large 494.7%estimated ± 1.1 pp, medium confidence
97Step 3.7 Flash94.7%estimated ± 1.1 pp, medium confidence
98GPT-5.4 nano94.7%estimated ± 1.1 pp, medium confidence
99o3-pro94.6%estimated ± 2.1 pp, medium confidence
100Grok 4.1 Fast (Reasoning)94.6%estimated ± 1.1 pp, medium confidence
101Mistral Large 394.6%estimated ± 1.1 pp, medium confidence
102Llama 4 Maverick94.5%estimated ± 1.1 pp, medium confidence
103Mistral Medium 3.5 128B94.5%estimated ± 1.1 pp, medium confidence
104Qwen3.6 Plus94.5%measured
105Qwen3.7 Plus94.5%measured
106Qwen3.5 397B (Reasoning)94.5%estimated ± 1.1 pp, medium confidence
107Qwen3.8-Flash-Next94.5%estimated ± 1.1 pp, medium confidence
108Qwen3 Max94.4%estimated ± 1.1 pp, medium confidence
109DeepSeek V3 032494.4%estimated ± 1.1 pp, medium confidence
110GLM-5.294.4%estimated ± 1.1 pp, medium confidence
111Nemotron 3 Super 100B94.4%estimated ± 1.1 pp, medium confidence
112DeepSeek V4 Pro 081394.4%estimated ± 0.7 pp, high confidence
113DeepSeek V3.294.3%estimated ± 1.1 pp, medium confidence
114GLM-5.194.3%estimated ± 1.1 pp, medium confidence
115Interfaze Beta94.2%estimated ± 2.1 pp, high confidence
116Grok Code Fast 194.2%estimated ± 1.1 pp, medium confidence
117Kimi K2.594.2%estimated ± 0.7 pp, high confidence
118Kimi K2.5 (Reasoning)94.2%estimated ± 0.7 pp, high confidence
119Llama 3.1 405B94.2%estimated ± 1.1 pp, medium confidence
120Pareto 26.10 Preview94.1%estimated ± 3.0 pp, medium confidence
121DeepSeek V3.194.1%estimated ± 1.1 pp, medium confidence
122Seed 2.1 Turbo94.1%estimated ± 0.9 pp, high confidence
123Grok 4 Fast (Reasoning)94.0%estimated ± 1.1 pp, medium confidence
124Claude 4 Sonnet94.0%estimated ± 1.1 pp, medium confidence
125Nemotron 3 Ultra94.0%estimated ± 0.7 pp, high confidence
126Trinity-Large-Preview94.0%estimated ± 1.1 pp, medium confidence
127Trinity-Large-Thinking94.0%estimated ± 1.1 pp, medium confidence
128Qwen3.5-122B-A10B93.9%estimated ± 0.7 pp, high confidence
129MiMo-V2.5-Pro93.9%estimated ± 1.1 pp, medium confidence
130Ornith-1.5-35B-A3B93.9%estimated ± 2.1 pp, medium confidence
131Mercury 2.593.8%estimated ± 1.1 pp, medium confidence
132GPT-OSS 120B93.8%estimated ± 1.1 pp, medium confidence
133Mistral Small 493.8%estimated ± 1.1 pp, medium confidence
134Mistral Small 4 (Reasoning)93.8%estimated ± 1.1 pp, medium confidence
135Solar Pro 493.7%estimated ± 0.7 pp, high confidence
136DeepSeek V4 Flash 073193.7%estimated ± 0.7 pp, high confidence
137Solar Open 293.7%estimated ± 0.7 pp, high confidence
138GLM-4.693.7%estimated ± 1.1 pp, medium confidence
139Claude 4.1 Opus93.7%estimated ± 2.1 pp, medium confidence
140Qwen3.5-27B93.6%estimated ± 0.7 pp, high confidence
141Beam93.5%estimated ± 3.0 pp, medium confidence
142Qwen3.6-27B93.5%measured
143GLM-593.4%estimated ± 0.7 pp, high confidence
144GPT-4.1 mini93.3%estimated ± 1.1 pp, medium confidence
145Nemotron Ultra 253B93.3%estimated ± 1.1 pp, medium confidence
146GPT-4o93.2%estimated ± 1.1 pp, medium confidence
147Mistral Large 293.2%estimated ± 1.1 pp, medium confidence
148Qwen3.5-35B-A3B93.2%estimated ± 0.7 pp, high confidence
149Gemma 4 31B93.1%estimated ± 0.7 pp, high confidence
150Qwen3.6-35B-A3B93.1%estimated ± 0.7 pp, high confidence
151MAI-Thinking-193.0%estimated ± 0.7 pp, high confidence
152Claude Sonnet 4.593.0%estimated ± 2.1 pp, high confidence
153MiMo-V2-Omni93.0%estimated ± 1.1 pp, medium confidence
154MiMo-V2-Flash93.0%estimated ± 0.7 pp, high confidence
155Ultravox v0.6 Llama 3.3 70B92.9%estimated ± 1.1 pp, medium confidence
156North Mini Code92.9%estimated ± 1.1 pp, medium confidence
157Grok 4.2092.8%estimated ± 3.0 pp, medium confidence
158A.X K292.7%estimated ± 1.1 pp, medium confidence
159Solar Pro 392.7%estimated ± 1.1 pp, medium confidence
160GLM-4.792.7%estimated ± 0.7 pp, high confidence
161Mistral Medium 392.6%estimated ± 1.1 pp, medium confidence
162Ling 3.0 Flash92.6%estimated ± 1.1 pp, medium confidence
163Ling 3.0 Flash FP892.6%estimated ± 1.1 pp, medium confidence
164Ornith-1.5-9B92.5%estimated ± 2.1 pp, medium confidence
165Claude 3 Haiku92.3%estimated ± 1.1 pp, medium confidence
166Sarvam 105B92.3%estimated ± 1.1 pp, medium confidence
167K-EXAONE 2.092.2%estimated ± 0.7 pp, high confidence
168Nemotron 3 Nano 30B92.2%estimated ± 1.1 pp, medium confidence
169Grok 4.1 Fast92.2%estimated ± 1.1 pp, medium confidence
170o1-pro92.1%estimated ± 2.1 pp, high confidence
171Nova Pro92.0%estimated ± 1.1 pp, medium confidence
172MiniMax M391.9%estimated ± 1.1 pp, medium confidence
173Qwen3 235B 250791.9%estimated ± 0.7 pp, high confidence
174K-Exaone91.8%estimated ± 1.1 pp, medium confidence
175GLM-4.5-Air91.8%estimated ± 1.1 pp, medium confidence
176Gemma 4 26B A4B91.7%estimated ± 0.7 pp, high confidence
177Solar Pro 291.7%estimated ± 1.1 pp, medium confidence
178o3-mini91.6%estimated ± 2.1 pp, high confidence
179GPT-OSS 20B91.6%estimated ± 1.1 pp, medium confidence
180Ling 2.6 Flash91.4%estimated ± 1.1 pp, medium confidence
181Qwen3.8-27B91.4%estimated ± 1.1 pp, medium confidence
182Claude Opus 4.691.4%estimated ± 0.7 pp, high confidence
183Quasar 438B91.4%estimated ± 1.1 pp, medium confidence
184Exaone 4.0 32B91.3%estimated ± 0.7 pp, high confidence
185Llama 4 Scout91.2%estimated ± 1.1 pp, medium confidence
186Nemotron 3.5 Lightning 30B A3B NVFP491.2%estimated ± 0.7 pp, high confidence
187Qwen3-Omni-30B-A3B-Thinking90.9%estimated ± 1.1 pp, medium confidence
188Ling 3.0 Flash VL90.7%estimated ± 1.1 pp, medium confidence
189Qwen3-Omni-30B-A3B-Instruct90.7%estimated ± 1.1 pp, medium confidence
190Phi-490.6%estimated ± 1.1 pp, medium confidence
191GPT-4.1 nano90.3%estimated ± 1.1 pp, medium confidence
192LongCat-Flash-Lite-Sparse89.8%estimated ± 0.7 pp, high confidence
193Claude Sonnet 4.689.8%estimated ± 0.7 pp, high confidence
194Gemma 3 27B89.8%estimated ± 1.1 pp, medium confidence
195Sarvam 30B89.5%estimated ± 1.1 pp, medium confidence
196Ternary Bonsai 2 27B89.1%measured
197Granite 4.2 30B88.9%estimated ± 0.7 pp, high confidence
198Nemotron 3 Nano Omni 30B A3B88.7%estimated ± 0.7 pp, high confidence
199Gemma 4 12B88.7%estimated ± 0.7 pp, high confidence
200Celeris-187.9%estimated ± 0.7 pp, high confidence
201DeepSeek V387.9%estimated ± 0.7 pp, high confidence
202ZAYA1-8B86.9%estimated ± 0.7 pp, high confidence
203Granite 4.2 8B86.8%estimated ± 0.7 pp, high confidence
204Claude 3.5 Sonnet86.6%estimated ± 2.1 pp, high confidence
205LFM2.5-8B-A1B86.3%estimated ± 1.1 pp, medium confidence
206Mellum2-12B-A2.5B-Thinking86.2%measured
207Command A+85.7%estimated ± 1.1 pp, medium confidence
208Ling 3.0 Tiny85.1%estimated ± 1.1 pp, medium confidence
209MiniCPM5-2B84.7%measured
210Gemma 4 E4B84.0%estimated ± 0.7 pp, high confidence
211ZAYA1-74B-Preview83.2%estimated ± 0.7 pp, high confidence
212Granite 4.2 3B83.0%estimated ± 0.7 pp, high confidence
213o1-preview81.3%estimated ± 2.1 pp, low confidence
214LFM2.5-VL-1.6B-Extract79.5%estimated ± 1.1 pp, low confidence
215Mellum2-12B-A2.5B-Instruct78.1%measured
216LLaDA2.2-mini78.1%estimated ± 3.0 pp, medium confidence
217Gemma 4 E2B77.9%estimated ± 0.7 pp, high confidence
218Granite-4.0-H-1B77.6%estimated ± 1.1 pp, low confidence
219Exaone 4.0 1.2B76.9%estimated ± 1.1 pp, low confidence
220LFM2.5-2.6B74.5%estimated ± 1.1 pp, low confidence
221Claude 3 Opus73.7%estimated ± 2.1 pp, low confidence
222DeepSeek R1 Distill Qwen 32B73.0%estimated ± 2.1 pp, low confidence
223Gemini 1.5 Pro72.2%estimated ± 2.1 pp, low confidence
224Granite-4.0-350M72.2%estimated ± 1.1 pp, low confidence
225Soofi S 30B-A3B71.9%estimated ± 0.7 pp, high confidence
226Granite-4.0-H-350M71.6%estimated ± 1.1 pp, low confidence
227GPT-4 Turbo71.2%estimated ± 2.1 pp, low confidence
228Qwen2.5 Coder 32B Instruct71.0%estimated ± 2.1 pp, low confidence
229GPT-4o mini70.9%estimated ± 2.1 pp, low confidence
230Phi-4 Multimodal Instruct70.5%estimated ± 2.1 pp, low confidence
231Gemini 1.0 Pro70.4%estimated ± 2.1 pp, low confidence
232MiniCPM5-1B70.1%measured
233LFM2.5-230M44.9%estimated ± 0.7 pp, medium confidence
234LFM2.5-VL-450M44.0%estimated ± 0.7 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General