benchgap
Knowledge & reasoning

GPQA Diamond leaderboard

As of 2026-10-07, the highest measured score on GPQA Diamond is 96.0% by GPT-6 Astra. 177 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 4 Argon97.1%estimated ± 2.2 pp, medium confidence
2Claude Sonnet 5.596.0%estimated ± 1.0 pp, low confidence
3Claude Opus 5.596.0%estimated ± 1.0 pp, low confidence
4Ling 3.1 Flash96.0%estimated ± 1.0 pp, low confidence
5GPT-6 Astra96.0%measured
6GPT-6.1 Sol96.0%estimated ± 1.0 pp, medium confidence
7Sakana Fugu95.5%measured
8Sakana Fugu-Ultra95.5%measured
9Claude Fable 595.2%estimated ± 1.4 pp, high confidence
10Gemini 3.8 Flash95.2%estimated ± 1.4 pp, high confidence
11Pareto 26.995.0%estimated ± 1.9 pp, high confidence
12GPT-6 Luna94.9%estimated ± 1.0 pp, medium confidence
13GPT-6 Sol94.9%estimated ± 1.0 pp, medium confidence
14Claude Fable 5.194.7%estimated ± 1.4 pp, high confidence
15MiMo-V2.6-Pro94.7%estimated ± 2.2 pp, high confidence
16GPT-5.6 Sol94.6%measured
17Muse Spark 1.394.5%estimated ± 2.2 pp, high confidence
18Gemini 3.1 Pro94.3%measured
19Claude Opus 4.7 (Adaptive)94.2%measured
20Claude Mythos 594.1%estimated ± 0.4 pp, high confidence
21Claude Opus 594.0%estimated ± 1.0 pp, medium confidence
22Claude Haiku 5.593.9%estimated ± 1.9 pp, high confidence
23Gemini 3.7 Flash93.8%estimated ± 1.4 pp, high confidence
24dots3-note Preview93.7%estimated ± 1.5 pp, high confidence
25Claude Opus 4.893.6%measured
26GPT-5.593.6%measured
27Muse Spark 1.193.6%estimated ± 1.0 pp, medium confidence
28GPT-5.5 Pro93.5%estimated ± 1.4 pp, high confidence
29Kimi K393.5%measured
30Step 5 Preview93.5%measured
31GPT-5.4 Pro93.3%estimated ± 1.4 pp, high confidence
32GPT-5.6 Terra92.9%measured
33GPT-5.492.8%measured
34Ornith-1.5-397B92.8%measured
35Gemini 3.5 Flash92.7%measured
36Grok 4.792.7%estimated ± 1.0 pp, medium confidence
37Qwen3.8 Max92.6%measured
38Claude Opus 4.6 (Adaptive)92.6%estimated ± 1.4 pp, high confidence
39Qwen3.8 Max Preview92.5%estimated ± 2.2 pp, high confidence
40Pareto 26.10 Preview92.4%measured
41Qwen3.7 Max92.4%measured
42GPT-5.292.4%estimated ± 0.4 pp, high confidence
43GPT-5.6 Luna92.3%measured
44Hy4 preview92.3%measured
45GPT-5.3 Codex92.3%estimated ± 2.2 pp, high confidence
46MiniMax M392.1%estimated ± 1.7 pp, high confidence
47Muse Spark 1.291.8%estimated ± 2.0 pp, high confidence
48Gemini 3.6 Flash91.8%estimated ± 1.4 pp, high confidence
49Qwen3.8-Flash-Next91.7%measured
50Agents-A191.4%estimated ± 2.3 pp, high confidence
51GLM-5.291.2%measured
52Qwen3.8-Omni-Flash91.0%measured
53DeepSeek V4.1 Flash90.9%measured
54Claude Opus 4.790.8%estimated ± 1.7 pp, high confidence
55Kimi K2.690.5%measured
56Beam90.5%measured
57Qwen3.6 Plus90.4%estimated ± 0.1 pp, high confidence
58Qwen3.7 Plus90.3%measured
59Grok 4.390.2%estimated ± 0.1 pp, high confidence
60DeepSeek V4 Pro 081390.1%measured
61Claude Sonnet 590.1%estimated ± 1.7 pp, high confidence
62Grok 4.690.0%estimated ± 1.4 pp, high confidence
63Interfaze Beta89.9%measured
64Claude Sonnet 4.689.9%estimated ± 0.1 pp, high confidence
65GLM-5.389.6%estimated ± 1.7 pp, high confidence
66Grok 4.589.5%estimated ± 1.4 pp, high confidence
67GPT-5.2-Codex89.5%estimated ± 2.2 pp, high confidence
68Gemini 3 Flash89.5%estimated ± 1.7 pp, high confidence
69Inkling-Small89.5%measured
70Muse Spark89.5%measured
71Gemini 3 Pro Deep Think89.4%estimated ± 1.5 pp, high confidence
72MiMo-V2.6-Flash89.3%estimated ± 2.2 pp, high confidence
73Kimi K2.7 Code89.2%estimated ± 2.2 pp, high confidence
74Mistral Large 489.2%estimated ± 2.2 pp, high confidence
75Claude Opus 4.689.2%measured
76Ornith-1.5-35B-A3B89.2%measured
77Qwen3.8-27B89.2%measured
78Solar Pro 489.0%measured
79Apodex 1.188.8%estimated ± 2.2 pp, high confidence
80Apodex 1.1 Mini88.8%estimated ± 2.2 pp, high confidence
81GLM-5.3-Flash88.7%estimated ± 1.7 pp, high confidence
82Hy388.5%estimated ± 2.2 pp, high confidence
83Grok 4.2088.5%measured
84Qwen3.5 397B88.4%estimated ± 0.4 pp, high confidence
85DeepSeek V4 Flash 073188.1%measured
86GPT-5.4 mini88.0%estimated ± 0.1 pp, high confidence
87Inkling87.9%measured
88Qwen3.6-27B87.8%estimated ± 0.4 pp, high confidence
89Kimi K2.587.6%measured
90Claude Opus 4.5 Thinking87.6%estimated ± 1.4 pp, medium confidence
91Kimi K2.5 (Reasoning)87.6%estimated ± 0.4 pp, high confidence
92Qwen 3.6 Max (preview)87.3%estimated ± 2.2 pp, high confidence
93Hy3 Preview87.2%measured
94Gemini 3.5 Flash-Lite87.2%estimated ± 1.7 pp, high confidence
95MiMo-V2-Pro87.1%estimated ± 2.2 pp, high confidence
96MiniMax M2.787.0%measured
97Nemotron 3 Ultra87.0%measured
98Claude Opus 4.587.0%estimated ± 0.4 pp, high confidence
99Qwen3.5-122B-A10B86.6%estimated ± 0.4 pp, high confidence
100Qwen3.5 Flash86.6%estimated ± 1.7 pp, medium confidence
101MiMo-V2.5-Pro86.4%estimated ± 1.7 pp, medium confidence
102Ornith-1.5-9B86.4%measured
103Solar Open 286.3%measured
104Gemini 3 Pro86.3%estimated ± 1.4 pp, medium confidence
105GLM-5.186.2%measured
106Seed 2.1 Pro86.2%estimated ± 12.9 pp, low confidence
107GPT-5 (high)86.1%estimated ± 2.2 pp, high confidence
108o3-pro86.0%estimated ± 2.4 pp, high confidence
109GLM-586.0%measured
110Qwen3.6-35B-A3B86.0%estimated ± 0.4 pp, high confidence
111MiMo-V2.585.8%estimated ± 1.7 pp, medium confidence
112GPT-5.185.8%estimated ± 1.4 pp, medium confidence
113Ternary Bonsai 2 27B85.8%measured
114GLM-5-Turbo85.8%estimated ± 2.2 pp, high confidence
115GLM-4.785.6%estimated ± 0.1 pp, high confidence
116A.X K285.6%measured
117Gemini 3.1 Flash-Lite85.5%estimated ± 1.7 pp, medium confidence
118Qwen3.5-27B85.5%estimated ± 0.4 pp, high confidence
119Grok 485.2%estimated ± 2.2 pp, high confidence
120Ling 3.0 Flash85.0%measured
121GPT-5.1-Codex84.6%estimated ± 2.2 pp, high confidence
122GPT-5.1-Codex-Max84.6%estimated ± 2.2 pp, high confidence
123GPT-5 (medium)84.5%estimated ± 2.2 pp, high confidence
124Claude Sonnet 4.5 Thinking84.3%estimated ± 1.4 pp, medium confidence
125Gemma 4 31B84.3%estimated ± 0.4 pp, high confidence
126MAI-Thinking-184.2%measured
127Qwen3.5-35B-A3B84.2%estimated ± 0.4 pp, high confidence
128Ling 3.0 Flash FP884.0%measured
129Seed 2.1 Turbo83.9%estimated ± 12.9 pp, low confidence
130Claude Sonnet 4.583.4%estimated ± 0.4 pp, high confidence
131MiMo-V2-Flash83.3%estimated ± 0.1 pp, high confidence
132Claude 4.1 Opus83.0%estimated ± 2.4 pp, high confidence
133Gemini 2.5 Pro83.0%estimated ± 0.4 pp, high confidence
134GPT-5.4 nano82.8%estimated ± 0.1 pp, high confidence
135MiMo-V2-Omni82.6%estimated ± 2.2 pp, high confidence
136Ling 3.0 Flash VL82.5%estimated ± 2.2 pp, high confidence
137Muse Glimmer 30B82.5%estimated ± 2.2 pp, high confidence
138K-EXAONE 2.082.2%measured
139Step 3.7 Flash82.2%estimated ± 2.2 pp, high confidence
140Nemotron 3 Super 100B81.8%estimated ± 2.2 pp, high confidence
141GLM-4.681.4%estimated ± 1.7 pp, medium confidence
142o381.4%estimated ± 2.2 pp, high confidence
143Qwen3.5 397B (Reasoning)81.2%estimated ± 2.2 pp, high confidence
144GPT-OSS 120B81.0%estimated ± 2.2 pp, high confidence
145Grok 4.1 Fast (Reasoning)80.9%estimated ± 2.2 pp, high confidence
146Grok 4 Fast (Reasoning)80.7%estimated ± 2.2 pp, high confidence
147Quasar 438B80.5%estimated ± 2.2 pp, high confidence
148Claude Haiku 4.579.9%estimated ± 1.7 pp, medium confidence
149GLM-4.579.9%estimated ± 1.7 pp, medium confidence
150Gemma 4 26B A4B79.7%estimated ± 1.9 pp, high confidence
151GLM-5V-Turbo79.4%estimated ± 2.2 pp, high confidence
152Mercury 2.579.0%measured
153o1-pro79.0%estimated ± 0.4 pp, high confidence
154Gemma 4 12B78.8%measured
155DeepSeek-R178.5%estimated ± 2.2 pp, high confidence
156Qwen3 235B 250777.5%estimated ± 0.4 pp, high confidence
157DeepSeek V3.1 (Reasoning)77.4%estimated ± 2.2 pp, high confidence
158o3-mini77.2%estimated ± 0.4 pp, high confidence
159K-Exaone77.1%estimated ± 2.2 pp, high confidence
160Trinity-Large-Thinking76.3%measured
161Claude 4.1 Opus Thinking76.1%estimated ± 2.2 pp, high confidence
162Command A+75.7%estimated ± 2.2 pp, high confidence
163o175.7%estimated ± 0.4 pp, high confidence
164Qwen3 Max75.6%estimated ± 2.2 pp, high confidence
165Nemotron 3.5 Lightning 30B A3B NVFP475.6%measured
166Nemotron 3 Nano 30B75.2%estimated ± 2.2 pp, high confidence
167DeepSeek V3.275.1%estimated ± 2.2 pp, high confidence
168North Mini Code75.0%estimated ± 2.2 pp, high confidence
169GPT-OSS 20B74.9%estimated ± 2.2 pp, high confidence
170Sarvam 105B74.9%estimated ± 2.2 pp, high confidence
171Solar Pro 374.3%estimated ± 2.2 pp, high confidence
172Mistral Small 474.0%estimated ± 2.2 pp, high confidence
173Mistral Small 4 (Reasoning)74.0%estimated ± 2.2 pp, high confidence
174Ling 3.0 Tiny73.5%estimated ± 2.2 pp, high confidence
175o1-preview72.6%estimated ± 2.4 pp, high confidence
176Grok Code Fast 172.4%estimated ± 2.2 pp, high confidence
177Nemotron 3 Nano Omni 30B A3B72.2%measured
178Qwen3-Omni-30B-A3B-Thinking71.9%estimated ± 2.2 pp, high confidence
179Sarvam 30B71.9%estimated ± 2.2 pp, high confidence
180Kimi K271.9%estimated ± 2.2 pp, high confidence
181Nemotron Ultra 253B71.9%estimated ± 2.2 pp, high confidence
182GLM-4.5-Air71.5%estimated ± 2.2 pp, high confidence
183LFM2.5-8B-A1B71.4%estimated ± 2.2 pp, high confidence
184Celeris-171.3%estimated ± 2.2 pp, high confidence
185DeepSeek V3.171.2%estimated ± 2.2 pp, high confidence
186ZAYA1-8B71.0%measured
187Granite-4.0-H-350M71.0%estimated ± 2.2 pp, high confidence
188LFM2.5-2.6B70.8%estimated ± 2.2 pp, high confidence
189Exaone 4.0 1.2B70.3%estimated ± 2.2 pp, high confidence
190MiniCPM5-2B70.2%measured
191Granite-4.0-350M70.1%estimated ± 2.2 pp, high confidence
192Grok 4.1 Fast69.7%estimated ± 2.2 pp, high confidence
193LFM2.5-VL-1.6B-Extract69.7%estimated ± 2.2 pp, high confidence
194Exaone 4.0 32B69.6%estimated ± 2.2 pp, high confidence
195Granite-4.0-H-1B69.6%estimated ± 2.2 pp, high confidence
196Phi-4 Multimodal Instruct69.6%estimated ± 2.2 pp, high confidence
197Llama 4 Maverick69.5%estimated ± 2.2 pp, high confidence
198LongCat-Flash-Lite-Sparse69.5%measured
199DeepSeek V3 032469.3%estimated ± 2.2 pp, medium confidence
200Gemini 2.5 Flash69.3%estimated ± 2.2 pp, medium confidence
201DeepSeek R1 Distill Qwen 32B69.2%estimated ± 2.2 pp, medium confidence
202Gemini 1.5 Pro69.2%estimated ± 2.2 pp, medium confidence
203Qwen3-Omni-30B-A3B-Instruct69.2%estimated ± 2.2 pp, medium confidence
204Gemma 3 27B69.0%estimated ± 2.2 pp, medium confidence
205Claude 4 Sonnet69.0%estimated ± 2.2 pp, medium confidence
206Gemini 1.0 Pro68.9%estimated ± 2.2 pp, medium confidence
207GPT-4o mini68.9%estimated ± 2.2 pp, medium confidence
208Mistral Large 368.9%estimated ± 2.2 pp, medium confidence
209Claude 3 Haiku68.8%estimated ± 2.2 pp, medium confidence
210Mistral Medium 368.8%estimated ± 2.2 pp, medium confidence
211Llama 3.1 405B68.7%estimated ± 2.2 pp, medium confidence
212Llama 4 Scout68.5%estimated ± 2.2 pp, medium confidence
213Phi-468.5%estimated ± 2.2 pp, medium confidence
214Solar Pro 268.4%estimated ± 2.2 pp, medium confidence
215Ultravox v0.6 Llama 3.3 70B68.3%estimated ± 2.2 pp, medium confidence
216Qwen2.5 Coder 32B Instruct68.2%estimated ± 2.2 pp, medium confidence
217Mistral Large 268.0%estimated ± 2.2 pp, medium confidence
218Nova Pro67.9%estimated ± 2.2 pp, medium confidence
219GPT-4 Turbo67.7%estimated ± 2.2 pp, medium confidence
220Claude 3 Opus67.4%estimated ± 2.2 pp, medium confidence
221Laguna XS.267.4%estimated ± 1.7 pp, medium confidence
222GPT-4o67.0%estimated ± 2.2 pp, medium confidence
223Granite 4.2 30B66.4%estimated ± 0.4 pp, high confidence
224GPT-4.166.3%estimated ± 0.4 pp, high confidence
225GPT-4.1 mini64.2%estimated ± 0.4 pp, high confidence
226Granite 4.2 8B64.1%estimated ± 0.4 pp, high confidence
227Trinity-Large-Preview63.3%measured
228Claude 3.5 Sonnet59.4%estimated ± 0.4 pp, high confidence
229DeepSeek V359.1%estimated ± 0.4 pp, high confidence
230Ling 2.6 Flash59.0%estimated ± 0.4 pp, high confidence
231Gemma 4 E4B58.6%estimated ± 0.4 pp, high confidence
232Mellum2-12B-A2.5B-Thinking57.6%measured
233ZAYA1-74B-Preview57.3%measured
234Granite 4.2 3B54.8%estimated ± 0.4 pp, high confidence
235GPT-4.1 nano50.3%estimated ± 0.4 pp, high confidence
236Mistral Medium 3.5 128B48.6%estimated ± 1.7 pp, medium confidence
237LLaDA2.2-mini44.4%measured
238Gemma 4 E2B43.4%estimated ± 0.4 pp, high confidence
239Soofi S 30B-A3B43.4%measured
240Mellum2-12B-A2.5B-Instruct40.9%measured
241Laguna M.139.9%estimated ± 1.7 pp, medium confidence
242MiniCPM5-1B26.3%measured
243LFM2.5-VL-450M25.7%estimated ± 0.4 pp, high confidence
244LFM2.5-230M25.4%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General