benchgap
Knowledge & reasoning

Vals GPQA Diamond leaderboard

As of 2026-10-07, the highest measured score on Vals GPQA Diamond is 95.5% by Gemini 3.1 Pro. 185 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.598.2%estimated ± 1.5 pp, low confidence
2Claude Opus 5.598.2%estimated ± 1.5 pp, low confidence
3Ling 3.1 Flash98.2%estimated ± 1.5 pp, low confidence
4GPT-6.1 Sol98.2%estimated ± 1.5 pp, low confidence
5Gemini 4 Argon96.4%estimated ± 2.2 pp, low confidence
6GPT-6 Luna96.2%estimated ± 1.5 pp, low confidence
7GPT-6 Sol96.2%estimated ± 1.5 pp, low confidence
8GPT-6 Astra96.1%estimated ± 1.1 pp, high confidence
9Gemini 3.1 Pro95.5%measured
10GPT-5.6 Sol95.2%measured
11Muse Spark 1.395.2%estimated ± 7.4 pp, medium confidence
12Sakana Fugu94.9%estimated ± 2.0 pp, medium confidence
13Sakana Fugu-Ultra94.9%estimated ± 2.0 pp, medium confidence
14Grok 4.694.7%measured
15Claude Mythos 594.4%estimated ± 3.0 pp, high confidence
16Gemini 3.8 Flash94.4%measured
17Gemini 3.7 Flash93.9%measured
18Qwen3.8 Max Preview93.8%estimated ± 7.4 pp, medium confidence
19Qwen3.8 Max93.7%measured
20GPT-5.5 Pro93.7%estimated ± 2.2 pp, high confidence
21GPT-5.4 Pro93.6%estimated ± 2.2 pp, high confidence
22Claude Fable 5.193.4%measured
23Claude Opus 593.4%measured
24Gemini 3.6 Flash93.4%measured
25dots3-note Preview93.4%estimated ± 2.2 pp, high confidence
26Claude Fable 593.2%measured
27GPT-5.593.2%measured
28Grok 4.592.9%measured
29Kimi K392.9%measured
30Step 5 Preview92.8%estimated ± 2.0 pp, high confidence
31Gemini 3.5 Flash92.7%measured
32MiniMax M392.7%measured
33Claude Opus 4.892.4%measured
34DeepSeek V4 Pro 081392.4%measured
35Pareto 26.992.3%estimated ± 3.0 pp, high confidence
36GPT-5.492.2%estimated ± 1.1 pp, high confidence
37Ornith-1.5-397B92.1%estimated ± 2.0 pp, high confidence
38Claude Opus 4.6 (Adaptive)92.1%estimated ± 2.2 pp, high confidence
39Pareto 26.10 Preview91.7%estimated ± 2.0 pp, high confidence
40GPT-5.6 Luna91.7%measured
41Claude Opus 4.7 (Adaptive)91.7%estimated ± 1.1 pp, high confidence
42Hy4 preview91.6%estimated ± 2.0 pp, high confidence
43Claude Haiku 5.591.5%estimated ± 3.0 pp, high confidence
44MiMo-V2.6-Pro91.4%estimated ± 7.5 pp, medium confidence
45GPT-5.3 Codex91.4%estimated ± 7.4 pp, medium confidence
46Grok 4.391.4%measured
47Muse Spark 1.191.2%measured
48Mistral Large 491.2%estimated ± 7.5 pp, medium confidence
49MiMo-V2.6-Flash91.1%estimated ± 7.5 pp, medium confidence
50Grok 4.791.1%estimated ± 1.5 pp, medium confidence
51Qwen3.8-Flash-Next91.0%estimated ± 2.0 pp, high confidence
52GPT-5.6 Terra90.9%measured
53Qwen3.8-Omni-Flash90.3%estimated ± 2.0 pp, high confidence
54Claude Opus 4.790.2%measured
55Qwen3.7 Max90.2%measured
56DeepSeek V4.1 Flash90.2%estimated ± 2.0 pp, high confidence
57GPT-5.290.0%estimated ± 2.2 pp, high confidence
58Muse Spark 1.290.0%estimated ± 3.8 pp, high confidence
59DeepSeek V4 Flash 073189.9%measured
60Beam89.8%estimated ± 2.0 pp, high confidence
61Muse Spark89.6%measured
62Qwen3.7 Plus89.6%estimated ± 2.0 pp, high confidence
63Interfaze Beta89.1%estimated ± 2.0 pp, high confidence
64Kimi K2.689.1%measured
65Claude Sonnet 588.9%measured
66Qwen3.8-27B88.9%measured
67Gemini 3 Pro Deep Think88.8%estimated ± 2.2 pp, high confidence
68Grok 4.2088.6%measured
69GPT-5.2-Codex88.5%estimated ± 7.4 pp, medium confidence
70Claude Opus 4.688.4%estimated ± 2.0 pp, high confidence
71Ornith-1.5-35B-A3B88.4%estimated ± 2.0 pp, high confidence
72Solar Pro 488.2%estimated ± 2.0 pp, high confidence
73Hy388.2%estimated ± 7.4 pp, medium confidence
74GLM-5.388.1%measured
75Kimi K2.7 Code88.0%estimated ± 7.4 pp, medium confidence
76Gemini 3 Flash87.9%measured
77Qwen3.6 Plus87.4%measured
78Claude Opus 4.5 Thinking87.3%estimated ± 2.2 pp, high confidence
79Inkling87.1%measured
80Kimi K2.586.8%estimated ± 2.0 pp, high confidence
81MiniMax M2.786.6%measured
82Qwen 3.6 Max (preview)86.6%estimated ± 7.4 pp, medium confidence
83GLM-5.3-Flash86.4%measured
84Hy3 Preview86.4%estimated ± 2.0 pp, high confidence
85Nemotron 3 Ultra86.1%measured
86Gemini 3 Pro85.9%estimated ± 2.2 pp, high confidence
87Claude Sonnet 4.685.6%measured
88GLM-5.285.6%measured
89Ornith-1.5-9B85.6%estimated ± 2.0 pp, high confidence
90Solar Open 285.4%estimated ± 2.0 pp, high confidence
91GPT-5.185.3%estimated ± 2.3 pp, high confidence
92GLM-585.1%estimated ± 2.0 pp, high confidence
93Ternary Bonsai 2 27B84.9%estimated ± 2.0 pp, high confidence
94Ling 3.0 Flash84.8%measured
95A.X K284.7%estimated ± 2.0 pp, high confidence
96Grok 484.7%estimated ± 7.4 pp, medium confidence
97Qwen3.5 397B84.6%estimated ± 4.8 pp, high confidence
98GLM-5.184.5%measured
99Gemini 3.5 Flash-Lite83.8%measured
100Inkling-Small83.6%measured
101Qwen3.6-27B83.5%estimated ± 4.8 pp, high confidence
102MiMo-V2-Pro83.5%estimated ± 7.4 pp, medium confidence
103MAI-Thinking-183.3%estimated ± 2.0 pp, medium confidence
104Kimi K2.5 (Reasoning)83.1%estimated ± 4.8 pp, high confidence
105GPT-5.4 mini83.1%measured
106Ling 3.0 Flash FP883.1%estimated ± 2.0 pp, medium confidence
107Qwen3.5 Flash82.8%measured
108MiMo-V2.5-Pro82.6%measured
109Apodex 1.182.4%estimated ± 7.4 pp, medium confidence
110Apodex 1.1 Mini82.4%estimated ± 7.4 pp, medium confidence
111Ling 3.0 Flash VL82.1%estimated ± 7.4 pp, medium confidence
112Claude Opus 4.582.1%estimated ± 4.8 pp, high confidence
113Qwen3.5 397B (Reasoning)81.9%estimated ± 7.4 pp, medium confidence
114GPT-5.1-Codex81.8%estimated ± 7.4 pp, medium confidence
115GPT-5.1-Codex-Max81.8%estimated ± 7.4 pp, medium confidence
116MiMo-V2.581.6%measured
117Qwen3.5-122B-A10B81.4%estimated ± 4.8 pp, high confidence
118K-EXAONE 2.081.2%estimated ± 2.0 pp, medium confidence
119Gemini 3.1 Flash-Lite81.1%measured
120GPT-5 (high)80.8%estimated ± 7.4 pp, medium confidence
121Claude Sonnet 4.5 Thinking80.7%estimated ± 2.2 pp, high confidence
122Claude Sonnet 4.580.7%estimated ± 2.2 pp, high confidence
123Grok 4.1 Fast (Reasoning)80.6%estimated ± 7.4 pp, medium confidence
124Qwen3.6-35B-A3B80.3%estimated ± 4.8 pp, high confidence
125GLM-4.780.1%measured
126GLM-5-Turbo79.6%estimated ± 7.4 pp, medium confidence
127Grok 4 Fast (Reasoning)79.6%estimated ± 7.4 pp, medium confidence
128Qwen3.5-27B79.5%estimated ± 4.8 pp, high confidence
129o3-pro79.3%estimated ± 7.4 pp, medium confidence
130GPT-5 (medium)78.8%estimated ± 7.4 pp, medium confidence
131Mercury 2.578.0%estimated ± 2.0 pp, medium confidence
132Gemma 4 12B77.7%estimated ± 2.0 pp, medium confidence
133Muse Glimmer 30B77.7%estimated ± 7.4 pp, medium confidence
134GPT-5.4 nano77.5%measured
135Qwen3.5-35B-A3B77.3%estimated ± 4.8 pp, high confidence
136Gemma 4 31B76.9%estimated ± 3.0 pp, medium confidence
137MiMo-V2-Omni76.6%estimated ± 7.4 pp, medium confidence
138Claude 4.1 Opus76.5%estimated ± 7.5 pp, medium confidence
139o376.4%estimated ± 7.4 pp, medium confidence
140Gemini 2.5 Pro75.3%estimated ± 4.8 pp, high confidence
141Trinity-Large-Thinking75.2%estimated ± 2.0 pp, medium confidence
142GLM-4.674.5%measured
143Nemotron 3.5 Lightning 30B A3B NVFP474.4%estimated ± 2.0 pp, medium confidence
144DeepSeek-R174.2%estimated ± 7.4 pp, medium confidence
145Claude 4.1 Opus Thinking73.6%estimated ± 7.4 pp, medium confidence
146GLM-5V-Turbo73.6%estimated ± 7.4 pp, medium confidence
147Step 3.7 Flash73.6%estimated ± 7.4 pp, medium confidence
148Nemotron 3 Super 100B72.3%estimated ± 7.4 pp, medium confidence
149Claude Haiku 4.572.2%measured
150GLM-4.572.2%measured
151Nemotron 3 Nano Omni 30B A3B71.0%estimated ± 2.0 pp, medium confidence
152K-Exaone69.7%estimated ± 7.4 pp, medium confidence
153ZAYA1-8B69.7%estimated ± 2.0 pp, medium confidence
154GPT-OSS 120B69.6%estimated ± 7.4 pp, medium confidence
155DeepSeek V3.1 (Reasoning)69.2%estimated ± 7.4 pp, medium confidence
156MiniCPM5-2B68.9%estimated ± 2.0 pp, medium confidence
157o1-pro68.8%estimated ± 4.8 pp, medium confidence
158LongCat-Flash-Lite-Sparse68.2%estimated ± 2.0 pp, medium confidence
159Mistral Small 467.7%estimated ± 7.4 pp, medium confidence
160Mistral Small 4 (Reasoning)67.7%estimated ± 7.4 pp, medium confidence
161Kimi K267.3%estimated ± 7.4 pp, medium confidence
162o1-preview67.1%estimated ± 7.4 pp, medium confidence
163Qwen3 Max67.0%estimated ± 7.4 pp, medium confidence
164Command A+66.6%estimated ± 7.4 pp, medium confidence
165Qwen3 235B 250766.5%estimated ± 4.8 pp, medium confidence
166o3-mini66.1%estimated ± 4.8 pp, medium confidence
167Nemotron 3 Nano 30B66.0%estimated ± 7.4 pp, medium confidence
168North Mini Code66.0%estimated ± 7.4 pp, medium confidence
169DeepSeek V3.265.2%estimated ± 7.4 pp, medium confidence
170o163.8%estimated ± 4.8 pp, medium confidence
171Sarvam 105B63.4%estimated ± 7.4 pp, medium confidence
172DeepSeek V3.163.0%estimated ± 7.4 pp, medium confidence
173Ling 3.0 Tiny62.9%estimated ± 7.4 pp, medium confidence
174GLM-4.5-Air62.7%estimated ± 7.4 pp, medium confidence
175Quasar 438B62.6%estimated ± 7.4 pp, medium confidence
176Nemotron Ultra 253B62.0%estimated ± 7.4 pp, medium confidence
177Grok Code Fast 161.9%estimated ± 7.4 pp, medium confidence
178Trinity-Large-Preview61.8%estimated ± 2.0 pp, medium confidence
179Qwen3-Omni-30B-A3B-Thinking61.8%estimated ± 7.4 pp, medium confidence
180Solar Pro 361.5%estimated ± 7.4 pp, medium confidence
181MiMo-V2-Flash59.3%measured
182Gemma 4 26B A4B57.2%estimated ± 3.0 pp, medium confidence
183GPT-OSS 20B56.8%estimated ± 7.4 pp, medium confidence
184Claude 4 Sonnet56.2%estimated ± 7.4 pp, medium confidence
185Gemini 2.5 Flash56.2%estimated ± 7.4 pp, medium confidence
186Mellum2-12B-A2.5B-Thinking56.0%estimated ± 2.0 pp, medium confidence
187Mistral Large 355.8%estimated ± 7.4 pp, medium confidence
188ZAYA1-74B-Preview55.7%estimated ± 2.0 pp, medium confidence
189Laguna XS.255.1%measured
190Llama 4 Maverick54.7%estimated ± 7.4 pp, medium confidence
191DeepSeek V3 032452.8%estimated ± 7.4 pp, medium confidence
192Granite 4.2 30B51.0%estimated ± 4.8 pp, medium confidence
193GPT-4.150.8%estimated ± 4.8 pp, medium confidence
194Grok 4.1 Fast50.7%estimated ± 7.4 pp, medium confidence
195Sarvam 30B50.2%estimated ± 7.4 pp, medium confidence
196Celeris-150.0%estimated ± 7.4 pp, low confidence
197Exaone 4.0 32B49.6%estimated ± 7.4 pp, low confidence
198Qwen3-Omni-30B-A3B-Instruct48.7%estimated ± 7.4 pp, low confidence
199DeepSeek R1 Distill Qwen 32B48.1%estimated ± 7.4 pp, low confidence
200GPT-4.1 mini48.1%estimated ± 4.8 pp, medium confidence
201Granite 4.2 8B48.1%estimated ± 4.8 pp, medium confidence
202Gemini 1.5 Pro45.3%estimated ± 7.4 pp, low confidence
203Llama 4 Scout45.0%estimated ± 7.4 pp, low confidence
204Mistral Medium 344.1%estimated ± 7.4 pp, low confidence
205Phi-443.7%estimated ± 7.4 pp, low confidence
206LLaDA2.2-mini42.4%estimated ± 2.0 pp, medium confidence
207Claude 3.5 Sonnet42.3%estimated ± 4.8 pp, medium confidence
208Solar Pro 242.3%estimated ± 7.4 pp, low confidence
209LFM2.5-2.6B41.9%estimated ± 7.4 pp, low confidence
210DeepSeek V341.9%estimated ± 4.8 pp, medium confidence
211Ling 2.6 Flash41.8%estimated ± 4.8 pp, medium confidence
212Soofi S 30B-A3B41.4%estimated ± 2.0 pp, medium confidence
213Gemma 4 E4B41.3%estimated ± 4.8 pp, medium confidence
214GPT-4o40.4%estimated ± 7.4 pp, low confidence
215Mellum2-12B-A2.5B-Instruct38.8%estimated ± 2.0 pp, medium confidence
216Llama 3.1 405B37.6%estimated ± 7.4 pp, low confidence
217LFM2.5-8B-A1B37.4%estimated ± 7.4 pp, low confidence
218Granite 4.2 3B37.0%estimated ± 4.8 pp, medium confidence
219Nova Pro36.0%estimated ± 7.4 pp, low confidence
220Ultravox v0.6 Llama 3.3 70B35.9%estimated ± 7.4 pp, low confidence
221Claude 3 Opus35.1%estimated ± 7.4 pp, low confidence
222Mistral Medium 3.5 128B34.8%measured
223Mistral Large 234.8%estimated ± 7.4 pp, low confidence
224GPT-4.1 nano32.1%estimated ± 4.8 pp, medium confidence
225Gemma 3 27B29.5%estimated ± 7.4 pp, low confidence
226GPT-4o mini29.3%estimated ± 7.4 pp, low confidence
227Exaone 4.0 1.2B29.2%estimated ± 7.4 pp, low confidence
228Qwen2.5 Coder 32B Instruct28.6%estimated ± 7.4 pp, low confidence
229Laguna M.127.0%measured
230Gemma 4 E2B25.1%estimated ± 4.8 pp, medium confidence
231Claude 3 Haiku24.9%estimated ± 7.4 pp, low confidence
232MiniCPM5-1B23.8%estimated ± 2.0 pp, medium confidence
233LFM2.5-230M22.9%estimated ± 2.0 pp, medium confidence
234Phi-4 Multimodal Instruct20.3%estimated ± 7.4 pp, low confidence
235LFM2.5-VL-1.6B-Extract18.3%estimated ± 7.4 pp, low confidence
236Gemini 1.0 Pro17.4%estimated ± 7.4 pp, low confidence
237Granite-4.0-H-1B16.4%estimated ± 7.4 pp, low confidence
238Granite-4.0-350M16.3%estimated ± 7.4 pp, low confidence
239Granite-4.0-H-350M16.0%estimated ± 7.4 pp, low confidence
240LFM2.5-VL-450M9.6%estimated ± 4.8 pp, medium confidence
241GPT-4 Turbo3.0%estimated ± 7.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General