benchgap
Knowledge & reasoning

Vals MMLU-Pro leaderboard

As of 2026-10-07, the highest measured score on Vals MMLU-Pro is 92.4% by Claude Fable 5.1. 186 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.596.7%estimated ± 1.5 pp, low confidence
2Claude Opus 5.594.0%estimated ± 1.5 pp, low confidence
3Ling 3.1 Flash93.8%estimated ± 1.5 pp, low confidence
4GPT-6.1 Sol93.0%estimated ± 1.5 pp, low confidence
5Gemini 4 Argon92.8%estimated ± 1.4 pp, low confidence
6Claude Fable 5.192.4%measured
7Claude Opus 591.6%measured
8Claude Fable 591.5%measured
9Gemini 3.1 Pro91.0%measured
10GPT-6 Luna90.3%estimated ± 1.5 pp, low confidence
11GPT-6 Sol90.3%estimated ± 1.5 pp, low confidence
12Gemini 3.8 Flash90.2%measured
13Gemini 3.7 Flash90.1%measured
14Claude Opus 4.789.9%measured
15dots3-note Preview89.9%estimated ± 2.0 pp, high confidence
16Pareto 26.989.7%estimated ± 1.9 pp, high confidence
17Claude Opus 4.889.6%measured
18Gemini 3.5 Flash89.5%measured
19Grok 4.689.4%measured
20GPT-5.5 Pro89.3%estimated ± 1.8 pp, high confidence
21Gemini 3.6 Flash89.3%measured
22Qwen3.7 Max89.3%measured
23GPT-5.4 Pro89.2%estimated ± 1.8 pp, high confidence
24GPT-5.3 Codex89.2%estimated ± 2.2 pp, high confidence
25Grok 4.589.2%measured
26Claude Haiku 5.589.1%estimated ± 1.9 pp, high confidence
27GPT-5.6 Sol89.1%measured
28Claude Opus 4.6 (Adaptive)89.0%estimated ± 1.8 pp, high confidence
29GPT-6 Astra88.9%estimated ± 1.2 pp, medium confidence
30Sakana Fugu88.8%estimated ± 1.2 pp, medium confidence
31Sakana Fugu-Ultra88.8%estimated ± 1.2 pp, medium confidence
32Muse Spark 1.188.7%measured
33Gemini 3 Flash88.6%measured
34Qwen3.8 Max88.6%measured
35Claude Opus 4.7 (Adaptive)88.5%estimated ± 1.2 pp, high confidence
36Step 5 Preview88.5%estimated ± 1.6 pp, high confidence
37Claude Mythos 588.5%estimated ± 1.2 pp, high confidence
38Muse Spark 1.288.3%measured
39GPT-5.488.1%estimated ± 1.2 pp, high confidence
40Ornith-1.5-397B88.1%estimated ± 1.2 pp, high confidence
41GPT-5.588.1%measured
42Kimi K388.0%measured
43GPT-5.288.0%estimated ± 1.2 pp, high confidence
44Hy4 preview87.9%estimated ± 1.2 pp, high confidence
45Pareto 26.10 Preview87.9%estimated ± 1.6 pp, high confidence
46Qwen3.8-Flash-Next87.7%estimated ± 1.2 pp, high confidence
47Qwen3.6 Plus87.7%measured
48Kimi K2.687.6%measured
49Claude Opus 4.687.5%estimated ± 1.2 pp, high confidence
50Muse Spark 1.387.5%estimated ± 2.2 pp, high confidence
51Claude Sonnet 587.5%measured
52Qwen3.8-Omni-Flash87.4%estimated ± 1.2 pp, high confidence
53DeepSeek V4.1 Flash87.4%estimated ± 1.2 pp, high confidence
54Claude Sonnet 4.687.3%measured
55Muse Spark87.3%measured
56Agents-A187.2%estimated ± 2.5 pp, medium confidence
57Qwen3.7 Plus87.0%estimated ± 1.2 pp, high confidence
58GPT-5.2-Codex87.0%estimated ± 2.2 pp, high confidence
59DeepSeek V4 Pro 081387.0%measured
60Grok 486.9%estimated ± 2.2 pp, high confidence
61GLM-5.186.9%measured
62GPT-5 (high)86.9%estimated ± 2.2 pp, high confidence
63Beam86.9%estimated ± 1.6 pp, high confidence
64Grok 4.786.8%estimated ± 1.5 pp, medium confidence
65GPT-5.1-Codex86.8%estimated ± 2.2 pp, high confidence
66GPT-5.1-Codex-Max86.8%estimated ± 2.2 pp, high confidence
67Interfaze Beta86.8%estimated ± 1.2 pp, high confidence
68GLM-5.386.8%measured
69Kimi K2.7 Code86.8%estimated ± 2.2 pp, high confidence
70GPT-5 (medium)86.7%estimated ± 2.2 pp, high confidence
71GLM-5.286.7%measured
72GPT-5.6 Terra86.7%measured
73o386.6%estimated ± 2.2 pp, high confidence
74Claude Opus 4.5 Thinking86.5%estimated ± 1.8 pp, high confidence
75Qwen 3.6 Max (preview)86.4%estimated ± 2.2 pp, high confidence
76Ornith-1.5-35B-A3B86.4%estimated ± 1.2 pp, high confidence
77Grok 4.2086.3%measured
78Inkling86.3%measured
79DeepSeek V4 Flash 073186.2%measured
80Gemini 3.1 Flash-Lite86.2%measured
81GLM-5.3-Flash86.1%measured
82Solar Pro 486.1%estimated ± 1.6 pp, high confidence
83GPT-5.6 Luna86.0%measured
84MiMo-V2.6-Pro85.9%estimated ± 2.2 pp, high confidence
85Gemini 3.5 Flash-Lite85.8%measured
86Grok 4.385.8%measured
87Nemotron 3 Ultra85.8%measured
88Qwen3.5 397B85.8%estimated ± 1.2 pp, high confidence
89Gemini 3 Pro Deep Think85.7%estimated ± 2.0 pp, high confidence
90Inkling-Small85.6%measured
91Gemini 3 Pro85.4%estimated ± 1.8 pp, high confidence
92Hy385.3%estimated ± 2.2 pp, high confidence
93Apodex 1.185.3%estimated ± 2.2 pp, high confidence
94Apodex 1.1 Mini85.3%estimated ± 2.2 pp, high confidence
95Qwen3.8 Max Preview85.3%estimated ± 2.2 pp, high confidence
96Qwen3.6-27B85.2%estimated ± 1.2 pp, high confidence
97DeepSeek-R185.1%estimated ± 2.2 pp, high confidence
98Kimi K2.585.0%estimated ± 1.2 pp, high confidence
99Kimi K2.5 (Reasoning)85.0%estimated ± 1.2 pp, high confidence
100GPT-5.184.9%estimated ± 1.8 pp, high confidence
101GLM-5V-Turbo84.8%estimated ± 2.2 pp, high confidence
102DeepSeek V3.1 (Reasoning)84.8%estimated ± 2.2 pp, high confidence
103GLM-5-Turbo84.7%estimated ± 2.2 pp, high confidence
104Hy3 Preview84.6%estimated ± 1.2 pp, high confidence
105GPT-5.4 mini84.6%measured
106MiMo-V2.5-Pro84.6%measured
107Solar Open 284.6%estimated ± 1.6 pp, high confidence
108Kimi K284.5%estimated ± 2.2 pp, high confidence
109Claude Opus 4.584.4%estimated ± 1.2 pp, high confidence
110MiMo-V2.6-Flash84.4%estimated ± 2.2 pp, high confidence
111Muse Glimmer 30B84.4%estimated ± 2.2 pp, high confidence
112MiMo-V2-Pro84.3%estimated ± 2.2 pp, high confidence
113Qwen3.8-27B84.3%measured
114Gemini 2.5 Flash84.3%estimated ± 2.2 pp, high confidence
115o3-pro84.2%estimated ± 2.6 pp, medium confidence
116MiniMax M384.2%measured
117Mistral Large 484.2%estimated ± 2.2 pp, high confidence
118Step 3.7 Flash84.2%estimated ± 2.2 pp, high confidence
119A.X K284.2%estimated ± 1.6 pp, high confidence
120Qwen3.5 Flash84.1%measured
121Grok 4.1 Fast (Reasoning)84.1%estimated ± 2.2 pp, high confidence
122Mistral Large 384.1%estimated ± 2.2 pp, high confidence
123Llama 4 Maverick84.0%estimated ± 2.2 pp, high confidence
124Qwen3.5-122B-A10B84.0%estimated ± 1.2 pp, high confidence
125Qwen3.5 397B (Reasoning)84.0%estimated ± 2.2 pp, high confidence
126Qwen3 Max83.9%estimated ± 2.2 pp, high confidence
127DeepSeek V3 032483.9%estimated ± 2.2 pp, high confidence
128Nemotron 3 Super 100B83.9%estimated ± 2.2 pp, high confidence
129DeepSeek V3.283.9%estimated ± 2.2 pp, high confidence
130Grok Code Fast 183.8%estimated ± 2.2 pp, high confidence
131Ornith-1.5-9B83.7%estimated ± 1.2 pp, high confidence
132Llama 3.1 405B83.7%estimated ± 2.2 pp, high confidence
133DeepSeek V3.183.7%estimated ± 2.2 pp, high confidence
134Grok 4 Fast (Reasoning)83.6%estimated ± 2.2 pp, high confidence
135Claude 4 Sonnet83.6%estimated ± 2.2 pp, high confidence
136GPT-OSS 120B83.5%estimated ± 2.2 pp, high confidence
137Mistral Small 483.4%estimated ± 2.2 pp, high confidence
138Mistral Small 4 (Reasoning)83.4%estimated ± 2.2 pp, high confidence
139GLM-583.2%estimated ± 1.2 pp, high confidence
140Qwen3.6-35B-A3B83.2%estimated ± 1.2 pp, high confidence
141Claude 4.1 Opus Thinking83.2%estimated ± 2.5 pp, high confidence
142Nemotron Ultra 253B83.1%estimated ± 2.2 pp, high confidence
143Phi-4 Multimodal Instruct83.1%estimated ± 2.5 pp, medium confidence
144DeepSeek R1 Distill Qwen 32B83.1%estimated ± 2.5 pp, medium confidence
145Gemini 1.5 Pro83.1%estimated ± 2.5 pp, medium confidence
146Gemini 1.0 Pro83.1%estimated ± 2.5 pp, medium confidence
147GPT-4o mini83.1%estimated ± 2.5 pp, medium confidence
148GPT-4o83.1%estimated ± 2.2 pp, high confidence
149Mistral Large 283.1%estimated ± 2.2 pp, high confidence
150Qwen2.5 Coder 32B Instruct83.1%estimated ± 2.5 pp, medium confidence
151GPT-4 Turbo83.1%estimated ± 2.5 pp, medium confidence
152Claude 3 Opus83.1%estimated ± 2.5 pp, medium confidence
153MiMo-V2-Omni83.0%estimated ± 2.2 pp, high confidence
154Ultravox v0.6 Llama 3.3 70B82.9%estimated ± 2.2 pp, high confidence
155North Mini Code82.9%estimated ± 2.2 pp, high confidence
156Ternary Bonsai 2 27B82.9%estimated ± 1.2 pp, high confidence
157MiMo-V2.582.9%measured
158Solar Pro 382.8%estimated ± 2.2 pp, high confidence
159Mistral Medium 382.8%estimated ± 2.2 pp, high confidence
160Claude 4.1 Opus82.7%estimated ± 2.6 pp, medium confidence
161GLM-4.782.7%measured
162Claude 3 Haiku82.7%estimated ± 2.2 pp, high confidence
163Sarvam 105B82.7%estimated ± 2.2 pp, high confidence
164Nemotron 3 Nano 30B82.6%estimated ± 2.2 pp, high confidence
165Grok 4.1 Fast82.6%estimated ± 2.2 pp, high confidence
166Nova Pro82.5%estimated ± 2.2 pp, high confidence
167Qwen3.5-27B82.5%estimated ± 1.2 pp, high confidence
168K-Exaone82.5%estimated ± 2.2 pp, high confidence
169GLM-4.5-Air82.4%estimated ± 2.2 pp, high confidence
170Solar Pro 282.4%estimated ± 2.2 pp, high confidence
171GPT-OSS 20B82.4%estimated ± 2.2 pp, high confidence
172Claude Sonnet 4.5 Thinking82.3%estimated ± 1.8 pp, high confidence
173Quasar 438B82.3%estimated ± 2.2 pp, medium confidence
174Llama 4 Scout82.2%estimated ± 2.2 pp, medium confidence
175GLM-4.682.2%measured
176K-EXAONE 2.082.2%estimated ± 1.6 pp, medium confidence
177Qwen3-Omni-30B-A3B-Thinking82.1%estimated ± 2.2 pp, medium confidence
178Ling 3.0 Flash VL82.1%estimated ± 2.2 pp, medium confidence
179Qwen3-Omni-30B-A3B-Instruct82.1%estimated ± 2.2 pp, medium confidence
180Phi-482.0%estimated ± 2.2 pp, medium confidence
181Ling 3.0 Flash82.0%measured
182Gemma 3 27B81.8%estimated ± 2.2 pp, medium confidence
183Sarvam 30B81.8%estimated ± 2.2 pp, medium confidence
184Celeris-181.5%estimated ± 2.2 pp, medium confidence
185Exaone 4.0 32B81.4%estimated ± 2.2 pp, medium confidence
186GLM-4.581.2%measured
187LFM2.5-8B-A1B81.2%estimated ± 2.2 pp, medium confidence
188Command A+81.1%estimated ± 2.2 pp, medium confidence
189Ling 3.0 Tiny81.0%estimated ± 2.2 pp, medium confidence
190Gemma 4 31B80.6%estimated ± 1.2 pp, high confidence
191LFM2.5-VL-1.6B-Extract80.5%estimated ± 2.2 pp, medium confidence
192MAI-Thinking-180.4%estimated ± 1.2 pp, high confidence
193Qwen3.5-35B-A3B80.4%estimated ± 1.2 pp, high confidence
194MiniMax M2.780.4%measured
195Granite-4.0-H-1B80.4%estimated ± 2.2 pp, medium confidence
196Exaone 4.0 1.2B80.3%estimated ± 2.2 pp, medium confidence
197Mercury 2.580.3%estimated ± 1.6 pp, medium confidence
198LFM2.5-2.6B80.2%estimated ± 2.2 pp, medium confidence
199Granite-4.0-350M80.1%estimated ± 2.2 pp, medium confidence
200Granite-4.0-H-350M80.1%estimated ± 2.2 pp, medium confidence
201Ling 3.0 Flash FP880.0%estimated ± 1.2 pp, high confidence
202MiMo-V2-Flash79.5%estimated ± 1.2 pp, high confidence
203Claude Sonnet 4.578.8%estimated ± 1.2 pp, high confidence
204Claude Haiku 4.578.7%measured
205Trinity-Large-Thinking78.6%estimated ± 1.6 pp, medium confidence
206Gemini 2.5 Pro78.0%estimated ± 1.2 pp, high confidence
207GPT-5.4 nano77.2%measured
208o1-preview77.0%estimated ± 2.6 pp, low confidence
209Mistral Medium 3.5 128B75.3%measured
210MiniCPM5-2B74.7%estimated ± 1.6 pp, medium confidence
211LongCat-Flash-Lite-Sparse74.2%estimated ± 1.6 pp, medium confidence
212Trinity-Large-Preview69.9%estimated ± 1.6 pp, medium confidence
213Laguna XS.269.1%measured
214Laguna M.168.8%measured
215o1-pro65.7%estimated ± 1.2 pp, medium confidence
216Gemma 4 12B64.9%estimated ± 1.2 pp, medium confidence
217Gemma 4 26B A4B61.2%estimated ± 1.9 pp, medium confidence
218Qwen3 235B 250759.4%estimated ± 1.2 pp, medium confidence
219o3-mini58.0%estimated ± 1.2 pp, medium confidence
220LLaDA2.2-mini54.8%estimated ± 1.6 pp, medium confidence
221o150.9%estimated ± 1.2 pp, medium confidence
222Nemotron 3.5 Lightning 30B A3B NVFP450.2%estimated ± 1.2 pp, medium confidence
223MiniCPM5-1B36.5%estimated ± 1.6 pp, medium confidence
224Nemotron 3 Nano Omni 30B A3B33.7%estimated ± 1.2 pp, medium confidence
225ZAYA1-8B28.2%estimated ± 1.2 pp, medium confidence
226Granite 4.2 30B12.7%estimated ± 1.2 pp, medium confidence
227GPT-4.112.4%estimated ± 1.2 pp, medium confidence
228GPT-4.1 mini8.2%estimated ± 1.2 pp, medium confidence
229Granite 4.2 8B8.1%estimated ± 1.2 pp, medium confidence
230Claude 3.5 Sonnet3.0%estimated ± 1.2 pp, medium confidence
231DeepSeek V32.8%estimated ± 1.2 pp, medium confidence
232Ling 2.6 Flash2.7%estimated ± 1.2 pp, medium confidence
233Gemma 4 E4B2.5%estimated ± 1.2 pp, medium confidence
234Mellum2-12B-A2.5B-Thinking2.0%estimated ± 1.2 pp, medium confidence
235ZAYA1-74B-Preview1.9%estimated ± 1.2 pp, medium confidence
236Granite 4.2 3B1.1%estimated ± 1.2 pp, medium confidence
237GPT-4.1 nano0.4%estimated ± 1.2 pp, medium confidence
238Gemma 4 E2B0.1%estimated ± 1.2 pp, medium confidence
239Soofi S 30B-A3B0.1%estimated ± 1.2 pp, medium confidence
240Mellum2-12B-A2.5B-Instruct0.0%estimated ± 1.2 pp, medium confidence
241LFM2.5-VL-450M0.0%estimated ± 1.2 pp, medium confidence
242LFM2.5-230M0.0%estimated ± 1.2 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General