benchgap
Knowledge & reasoning

GPQA leaderboard

As of 2026-10-07, the highest measured score on GPQA is 96.0% by GPT-6 Astra. 160 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.596.0%estimated ± 1.0 pp, low confidence
2Claude Opus 5.596.0%estimated ± 1.0 pp, low confidence
3Ling 3.1 Flash96.0%estimated ± 1.0 pp, low confidence
4GPT-6 Astra96.0%measured
5GPT-6.1 Sol95.9%estimated ± 0.8 pp, high confidence
6Claude Fable 5.195.9%estimated ± 1.6 pp, medium confidence
7Claude Opus 595.6%estimated ± 0.8 pp, high confidence
8Sakana Fugu95.5%measured
9Sakana Fugu-Ultra95.5%measured
10Claude Fable 595.1%estimated ± 1.6 pp, medium confidence
11Gemini 3.8 Flash94.8%estimated ± 0.8 pp, high confidence
12Pareto 26.994.8%estimated ± 2.1 pp, high confidence
13GPT-5.6 Sol94.6%measured
14Gemini 3.1 Pro94.2%estimated ± 0.1 pp, high confidence
15Claude Opus 4.7 (Adaptive)94.2%measured
16GPT-6 Sol94.2%estimated ± 0.8 pp, high confidence
17Claude Mythos 594.1%measured
18GPT-5.5 Pro93.8%estimated ± 1.6 pp, high confidence
19Gemini 3.7 Flash93.7%estimated ± 1.6 pp, medium confidence
20GPT-5.4 Pro93.7%estimated ± 1.6 pp, high confidence
21Grok 4.693.7%estimated ± 0.8 pp, high confidence
22Claude Haiku 5.593.6%estimated ± 2.1 pp, high confidence
23Claude Opus 4.893.6%measured
24GPT-5.593.6%measured
25Muse Spark 1.193.6%estimated ± 1.0 pp, medium confidence
26dots3-note Preview93.6%estimated ± 1.6 pp, high confidence
27Claude Opus 4.793.5%estimated ± 1.6 pp, medium confidence
28Step 5 Preview93.5%estimated ± 0.4 pp, high confidence
29Kimi K393.5%measured
30Grok 4.593.2%estimated ± 0.8 pp, high confidence
31GPT-6 Luna93.1%estimated ± 0.8 pp, medium confidence
32Gemini 3.6 Flash93.0%estimated ± 1.6 pp, high confidence
33GPT-5.6 Terra92.9%measured
34Gemini 4 Argon92.8%estimated ± 3.2 pp, high confidence
35GPT-5.492.8%measured
36Ornith-1.5-397B92.8%measured
37Grok 4.792.7%estimated ± 1.0 pp, medium confidence
38Qwen3.8 Max92.6%measured
39Muse Spark 1.392.6%estimated ± 3.2 pp, high confidence
40MiMo-V2.6-Pro92.5%estimated ± 3.2 pp, high confidence
41Claude Opus 4.6 (Adaptive)92.4%estimated ± 1.6 pp, high confidence
42Pareto 26.10 Preview92.4%estimated ± 0.4 pp, high confidence
43Qwen3.8 Max Preview92.4%estimated ± 3.2 pp, high confidence
44GPT-5.292.4%measured
45Qwen3.7 Max92.4%measured
46Gemini 3 Flash92.3%estimated ± 1.6 pp, high confidence
47GPT-5.6 Luna92.3%measured
48Hy4 preview92.3%measured
49Gemini 3.5 Flash92.2%measured
50Apodex 1.192.1%estimated ± 2.2 pp, high confidence
51Muse Spark 1.292.0%estimated ± 1.6 pp, high confidence
52Mistral Large 491.8%estimated ± 3.2 pp, high confidence
53Qwen3.8-Flash-Next91.7%measured
54MiMo-V2.6-Flash91.7%estimated ± 3.2 pp, high confidence
55Claude Opus 4.691.3%measured
56Agents-A191.3%estimated ± 2.2 pp, high confidence
57Claude Sonnet 591.3%estimated ± 1.6 pp, high confidence
58GLM-5.291.2%measured
59Qwen3.8-Omni-Flash91.0%measured
60DeepSeek V4.1 Flash90.9%measured
61GPT-5.3 Codex90.8%estimated ± 3.2 pp, high confidence
62GLM-5.390.6%estimated ± 1.6 pp, high confidence
63Beam90.5%estimated ± 0.4 pp, high confidence
64Kimi K2.690.5%measured
65Qwen3.6 Plus90.4%measured
66Qwen3.7 Plus90.3%measured
67DeepSeek V4 Pro 081390.1%measured
68Grok 4.390.1%measured
69Gemini 3.1 Flash-Lite90.0%estimated ± 1.6 pp, high confidence
70GLM-5.3-Flash89.9%estimated ± 1.6 pp, high confidence
71Claude Sonnet 4.689.9%measured
72Interfaze Beta89.9%measured
73Gemini 3 Pro Deep Think89.9%estimated ± 1.6 pp, high confidence
74MiMo-V2-Pro89.9%estimated ± 3.2 pp, high confidence
75GPT-5.2-Codex89.8%estimated ± 3.2 pp, high confidence
76Gemini 3.5 Flash-Lite89.6%estimated ± 1.6 pp, high confidence
77Muse Spark89.5%estimated ± 0.1 pp, high confidence
78Inkling-Small89.5%measured
79Qwen 3.6 Max (preview)89.4%estimated ± 2.6 pp, high confidence
80Quasar 438B89.2%estimated ± 3.2 pp, high confidence
81Ornith-1.5-35B-A3B89.2%measured
82Qwen3.8-27B89.2%measured
83GLM-5-Turbo89.2%estimated ± 3.2 pp, high confidence
84Apodex 1.1 Mini89.1%estimated ± 3.2 pp, high confidence
85Solar Pro 489.0%estimated ± 0.4 pp, high confidence
86Claude Opus 4.5 Thinking88.9%estimated ± 1.6 pp, high confidence
87Kimi K2.7 Code88.8%estimated ± 3.2 pp, high confidence
88Hy388.6%estimated ± 3.2 pp, high confidence
89Grok 4.2088.5%estimated ± 0.1 pp, high confidence
90MiMo-V2.5-Pro88.5%estimated ± 1.6 pp, high confidence
91Qwen3.5 397B88.4%measured
92Ling 3.0 Flash VL88.3%estimated ± 3.2 pp, high confidence
93Seed 2.1 Pro88.3%estimated ± 2.6 pp, high confidence
94MiniMax M388.1%estimated ± 1.6 pp, high confidence
95DeepSeek V4 Flash 073188.1%measured
96GPT-5.188.1%estimated ± 1.7 pp, high confidence
97Qwen3.5 Flash88.0%estimated ± 1.6 pp, high confidence
98GPT-5.4 mini88.0%measured
99Gemini 3 Pro88.0%estimated ± 1.6 pp, high confidence
100MiMo-V2-Omni88.0%estimated ± 3.2 pp, high confidence
101Inkling87.9%measured
102GPT-5.1-Codex87.8%estimated ± 3.2 pp, high confidence
103GPT-5.1-Codex-Max87.8%estimated ± 3.2 pp, high confidence
104Qwen3.6-27B87.8%measured
105GLM-5V-Turbo87.7%estimated ± 3.2 pp, high confidence
106Kimi K2.587.6%measured
107Kimi K2.5 (Reasoning)87.6%measured
108GPT-5 (high)87.4%estimated ± 3.2 pp, high confidence
109GPT-5 (medium)87.4%estimated ± 3.2 pp, high confidence
110Claude 4.1 Opus Thinking87.3%estimated ± 3.2 pp, high confidence
111Hy3 Preview87.2%measured
112Command A+87.1%estimated ± 3.2 pp, high confidence
113Grok 487.1%estimated ± 3.2 pp, high confidence
114MiniMax M2.787.0%estimated ± 0.1 pp, high confidence
115Claude Opus 4.587.0%measured
116Nemotron 3 Ultra87.0%measured
117MiMo-V2.586.9%estimated ± 1.6 pp, high confidence
118o3-pro86.7%estimated ± 3.2 pp, high confidence
119Qwen3.5-122B-A10B86.6%measured
120Seed 2.1 Turbo86.6%estimated ± 2.6 pp, high confidence
121Qwen3.5 397B (Reasoning)86.4%estimated ± 3.2 pp, high confidence
122Ornith-1.5-9B86.4%measured
123Solar Open 286.3%estimated ± 0.4 pp, high confidence
124GLM-5.186.2%estimated ± 0.1 pp, high confidence
125GLM-4.686.2%estimated ± 1.6 pp, high confidence
126GLM-586.0%measured
127Qwen3.6-35B-A3B86.0%measured
128Ternary Bonsai 2 27B85.8%measured
129GLM-4.785.7%measured
130A.X K285.6%estimated ± 0.4 pp, high confidence
131Grok 4.1 Fast (Reasoning)85.6%estimated ± 3.2 pp, high confidence
132Qwen3.5-27B85.5%measured
133o385.4%estimated ± 3.2 pp, high confidence
134GLM-4.585.2%estimated ± 1.6 pp, high confidence
135Ling 3.0 Flash85.0%measured
136Claude Sonnet 4.5 Thinking84.8%estimated ± 1.6 pp, high confidence
137Step 3.7 Flash84.8%estimated ± 3.2 pp, high confidence
138Gemma 4 31B84.3%measured
139MAI-Thinking-184.2%measured
140Qwen3.5-35B-A3B84.2%measured
141Ling 3.0 Flash FP884.0%measured
142Claude 4.1 Opus83.8%estimated ± 3.2 pp, high confidence
143MiMo-V2-Flash83.7%measured
144Claude Sonnet 4.583.4%measured
145Grok 4 Fast (Reasoning)83.1%estimated ± 3.2 pp, high confidence
146Gemini 2.5 Pro83.0%measured
147Claude Haiku 4.582.9%estimated ± 1.6 pp, high confidence
148GPT-5.4 nano82.8%measured
149Muse Glimmer 30B82.5%estimated ± 3.2 pp, high confidence
150K-EXAONE 2.082.2%estimated ± 0.4 pp, high confidence
151Claude 4 Sonnet81.3%estimated ± 3.2 pp, high confidence
152DeepSeek V3.280.4%estimated ± 3.2 pp, high confidence
153Gemma 4 26B A4B80.2%estimated ± 2.1 pp, high confidence
154Qwen3 Max79.7%estimated ± 3.2 pp, high confidence
155Mistral Medium 3.5 128B79.6%estimated ± 1.6 pp, medium confidence
156Mercury 2.579.0%estimated ± 0.4 pp, high confidence
157o1-pro79.0%measured
158Gemma 4 12B78.8%measured
159Qwen3 235B 250777.5%measured
160K-Exaone77.2%estimated ± 3.2 pp, high confidence
161o3-mini77.2%measured
162Grok Code Fast 176.5%estimated ± 3.2 pp, high confidence
163Trinity-Large-Thinking76.3%estimated ± 0.4 pp, high confidence
164o175.7%measured
165DeepSeek V3.175.7%estimated ± 3.2 pp, high confidence
166Nemotron 3.5 Lightning 30B A3B NVFP475.6%measured
167DeepSeek V3.1 (Reasoning)75.1%estimated ± 3.2 pp, high confidence
168DeepSeek-R174.2%estimated ± 3.2 pp, high confidence
169Laguna XS.273.7%estimated ± 1.6 pp, medium confidence
170Laguna M.173.4%estimated ± 1.6 pp, medium confidence
171Nemotron 3 Super 100B73.3%estimated ± 3.2 pp, high confidence
172Kimi K273.0%estimated ± 3.2 pp, high confidence
173Nemotron 3 Nano Omni 30B A3B72.2%measured
174ZAYA1-8B71.0%measured
175MiniCPM5-2B70.2%estimated ± 0.4 pp, high confidence
176LongCat-Flash-Lite-Sparse69.5%estimated ± 0.4 pp, high confidence
177GPT-OSS 120B69.4%estimated ± 3.2 pp, high confidence
178o1-preview68.6%estimated ± 3.2 pp, high confidence
179Grok 4.1 Fast68.2%estimated ± 3.2 pp, high confidence
180Mistral Small 468.2%estimated ± 3.2 pp, high confidence
181Mistral Small 4 (Reasoning)68.2%estimated ± 3.2 pp, high confidence
182GLM-4.5-Air67.5%estimated ± 3.2 pp, high confidence
183Ling 3.0 Tiny67.4%estimated ± 3.2 pp, high confidence
184Granite 4.2 30B66.4%measured
185GPT-4.166.3%measured
186GPT-4.1 mini64.2%measured
187Granite 4.2 8B64.1%measured
188Trinity-Large-Preview63.3%estimated ± 0.4 pp, high confidence
189Llama 4 Maverick62.7%estimated ± 3.2 pp, high confidence
190North Mini Code62.3%estimated ± 3.2 pp, high confidence
191Gemini 2.5 Flash62.0%estimated ± 3.2 pp, high confidence
192DeepSeek V3 032461.4%estimated ± 3.2 pp, high confidence
193Claude 3.5 Sonnet59.4%measured
194DeepSeek V359.1%measured
195Mistral Large 359.0%estimated ± 3.2 pp, high confidence
196Ling 2.6 Flash59.0%measured
197Gemma 4 E4B58.6%measured
198Mistral Medium 357.8%estimated ± 3.2 pp, high confidence
199Mellum2-12B-A2.5B-Thinking57.6%measured
200GPT-OSS 20B57.4%estimated ± 3.2 pp, high confidence
201ZAYA1-74B-Preview57.3%measured
202Nemotron 3 Nano 30B57.0%estimated ± 3.2 pp, high confidence
203Sarvam 105B56.3%estimated ± 3.2 pp, high confidence
204Claude 3 Opus55.9%estimated ± 3.2 pp, high confidence
205Granite 4.2 3B54.8%measured
206GPT-4o54.2%estimated ± 3.2 pp, high confidence
207LFM2.5-2.6B53.9%estimated ± 3.2 pp, high confidence
208DeepSeek R1 Distill Qwen 32B53.8%estimated ± 3.2 pp, high confidence
209Llama 4 Scout51.9%estimated ± 3.2 pp, high confidence
210Gemini 1.5 Pro50.7%estimated ± 3.2 pp, high confidence
211GPT-4.1 nano50.3%measured
212Solar Pro 350.1%estimated ± 3.2 pp, high confidence
213Qwen3-Omni-30B-A3B-Thinking49.7%estimated ± 3.2 pp, medium confidence
214Ultravox v0.6 Llama 3.3 70B49.0%estimated ± 3.2 pp, medium confidence
215Mistral Large 248.3%estimated ± 3.2 pp, medium confidence
216Nemotron Ultra 253B48.1%estimated ± 3.2 pp, medium confidence
217Llama 3.1 405B46.4%estimated ± 3.2 pp, medium confidence
218LFM2.5-8B-A1B45.8%estimated ± 3.2 pp, medium confidence
219GPT-4 Turbo44.5%estimated ± 3.2 pp, medium confidence
220LLaDA2.2-mini44.4%estimated ± 0.4 pp, high confidence
221Solar Pro 244.2%estimated ± 3.2 pp, medium confidence
222Nova Pro43.9%estimated ± 3.2 pp, medium confidence
223Gemma 4 E2B43.4%measured
224Soofi S 30B-A3B43.4%measured
225Qwen2.5 Coder 32B Instruct42.1%estimated ± 3.2 pp, medium confidence
226GPT-4o mini41.5%estimated ± 3.2 pp, medium confidence
227Mellum2-12B-A2.5B-Instruct40.9%measured
228Sarvam 30B40.7%estimated ± 3.2 pp, medium confidence
229Celeris-139.0%estimated ± 3.2 pp, medium confidence
230Exaone 4.0 32B38.8%estimated ± 3.2 pp, medium confidence
231Qwen3-Omni-30B-A3B-Instruct36.1%estimated ± 3.2 pp, medium confidence
232Phi-435.4%estimated ± 3.2 pp, medium confidence
233Phi-4 Multimodal Instruct34.4%estimated ± 3.2 pp, medium confidence
234Claude 3 Haiku32.4%estimated ± 3.2 pp, medium confidence
235Gemini 1.0 Pro30.3%estimated ± 3.2 pp, medium confidence
236Exaone 4.0 1.2B29.3%estimated ± 3.2 pp, medium confidence
237Granite-4.0-H-1B29.0%estimated ± 3.2 pp, medium confidence
238MiniCPM5-1B26.3%estimated ± 0.4 pp, high confidence
239Gemma 3 27B26.0%estimated ± 3.2 pp, medium confidence
240Granite-4.0-350M25.8%estimated ± 3.2 pp, medium confidence
241Granite-4.0-H-350M25.8%estimated ± 3.2 pp, medium confidence
242LFM2.5-VL-1.6B-Extract25.8%estimated ± 3.2 pp, medium confidence
243LFM2.5-VL-450M25.7%measured
244LFM2.5-230M25.4%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General