benchgap
Knowledge & reasoning

SuperGPQA leaderboard

As of 2026-10-07, the highest measured score on SuperGPQA is 95.0% by Claude Opus 4.6. 214 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 4.695.0%measured
2Claude Sonnet 4.695.0%measured
3Claude Haiku 5.591.4%estimated ± 10.4 pp, low confidence
4Claude Fable 5.190.9%estimated ± 7.5 pp, low confidence
5Claude Opus 5.590.7%estimated ± 7.5 pp, low confidence
6Claude Fable 590.5%estimated ± 7.5 pp, low confidence
7GLM-5.3-Flash90.3%estimated ± 10.4 pp, low confidence
8GPT-6 Astra89.7%estimated ± 7.5 pp, low confidence
9GPT-6.1 Sol89.6%estimated ± 7.5 pp, low confidence
10Claude Opus 589.3%estimated ± 7.5 pp, low confidence
11GPT-5.6 Sol88.8%estimated ± 7.5 pp, low confidence
12GPT-5.588.4%estimated ± 7.5 pp, low confidence
13Gemini 3 Pro87.7%estimated ± 7.5 pp, low confidence
14Claude Mythos 587.6%estimated ± 9.1 pp, low confidence
15Gemini 3.7 Flash87.6%estimated ± 7.5 pp, low confidence
16Gemini 3.1 Pro87.4%estimated ± 7.5 pp, low confidence
17Gemini 3.8 Flash87.3%estimated ± 7.5 pp, low confidence
18GPT-6 Sol87.3%estimated ± 7.5 pp, low confidence
19Claude Sonnet 5.587.1%estimated ± 7.5 pp, low confidence
20GPT-5.3 Codex86.8%estimated ± 7.5 pp, low confidence
21Muse Spark 1.186.5%estimated ± 7.5 pp, low confidence
22Gemini 3.5 Flash86.4%estimated ± 7.5 pp, low confidence
23Grok 4.586.3%estimated ± 7.5 pp, low confidence
24GPT-5.486.0%estimated ± 7.5 pp, low confidence
25GPT-5.4 Pro85.9%estimated ± 9.1 pp, low confidence
26Gemini 3.6 Flash85.7%estimated ± 7.5 pp, low confidence
27Gemini 4 Argon85.7%estimated ± 7.5 pp, low confidence
28Muse Spark85.5%estimated ± 7.5 pp, low confidence
29GPT-5.5 Pro85.4%estimated ± 9.1 pp, low confidence
30DeepSeek V4 Pro 081385.3%estimated ± 7.5 pp, low confidence
31Claude Opus 4.7 (Adaptive)85.3%estimated ± 7.5 pp, low confidence
32Claude Opus 4.885.2%estimated ± 7.5 pp, low confidence
33Grok 4.685.0%estimated ± 7.5 pp, low confidence
34Hy4 preview84.8%estimated ± 9.1 pp, low confidence
35Kimi K384.7%estimated ± 7.5 pp, low confidence
36Grok 4.784.7%estimated ± 7.5 pp, low confidence
37Claude Opus 4.6 (Adaptive)84.5%estimated ± 7.5 pp, low confidence
38GPT-5.6 Terra84.4%estimated ± 7.5 pp, low confidence
39Claude Opus 4.5 Thinking84.3%estimated ± 7.5 pp, low confidence
40DeepSeek V4.1 Flash84.2%estimated ± 7.5 pp, low confidence
41Gemini 3 Flash84.0%estimated ± 7.5 pp, medium confidence
42dots3-note Preview83.8%estimated ± 9.1 pp, medium confidence
43Muse Spark 1.283.8%estimated ± 7.5 pp, medium confidence
44Claude Opus 4.783.5%estimated ± 7.5 pp, medium confidence
45Sakana Fugu83.5%estimated ± 9.2 pp, low confidence
46Sakana Fugu-Ultra83.5%estimated ± 9.2 pp, low confidence
47GPT-5.283.3%estimated ± 7.5 pp, medium confidence
48GPT-6 Luna83.1%estimated ± 7.5 pp, medium confidence
49Muse Spark 1.383.0%estimated ± 7.5 pp, medium confidence
50GPT-5.6 Luna82.5%estimated ± 7.5 pp, medium confidence
51Inkling82.0%estimated ± 7.5 pp, medium confidence
52Step 5 Preview81.9%estimated ± 7.5 pp, medium confidence
53Agents-A181.8%estimated ± 9.1 pp, medium confidence
54GPT-5.2-Codex81.7%estimated ± 7.5 pp, medium confidence
55Grok 481.4%estimated ± 7.5 pp, medium confidence
56DeepSeek V4 Flash 073181.4%estimated ± 7.5 pp, medium confidence
57GPT-5 (high)81.3%estimated ± 7.5 pp, medium confidence
58Claude Sonnet 581.2%estimated ± 7.5 pp, medium confidence
59GPT-5.1-Codex81.1%estimated ± 7.5 pp, medium confidence
60GPT-5.1-Codex-Max81.1%estimated ± 7.5 pp, medium confidence
61Kimi K2.7 Code80.9%estimated ± 7.5 pp, medium confidence
62GPT-5 (medium)80.9%estimated ± 7.5 pp, medium confidence
63Gemini 2.5 Pro80.7%estimated ± 7.5 pp, medium confidence
64Ornith-1.5-397B80.5%estimated ± 9.1 pp, medium confidence
65o380.4%estimated ± 7.5 pp, medium confidence
66Pareto 26.10 Preview80.2%estimated ± 13.9 pp, low confidence
67Qwen3.8 Max80.0%estimated ± 9.1 pp, medium confidence
68GPT-5.179.9%estimated ± 7.5 pp, medium confidence
69GPT-5.4 mini79.7%estimated ± 7.5 pp, medium confidence
70Kimi K2.5 (Reasoning)78.3%estimated ± 7.5 pp, medium confidence
71MiMo-V2.6-Pro78.1%estimated ± 7.5 pp, medium confidence
72Grok 4.377.9%estimated ± 7.5 pp, medium confidence
73o177.9%estimated ± 7.5 pp, medium confidence
74GLM-5.377.5%estimated ± 7.5 pp, medium confidence
75Beam77.2%estimated ± 13.9 pp, low confidence
76Inkling-Small77.0%estimated ± 7.5 pp, medium confidence
77Kimi K2.676.6%estimated ± 7.5 pp, medium confidence
78Hy376.1%estimated ± 7.5 pp, medium confidence
79Qwen3.8-Omni-Flash76.0%estimated ± 9.1 pp, medium confidence
80Apodex 1.175.9%estimated ± 7.5 pp, medium confidence
81Apodex 1.1 Mini75.9%estimated ± 7.5 pp, medium confidence
82Qwen3.8 Max Preview75.9%estimated ± 7.5 pp, medium confidence
83Hy3 Preview75.7%estimated ± 7.5 pp, medium confidence
84Interfaze Beta75.4%estimated ± 9.2 pp, low confidence
85DeepSeek-R175.0%estimated ± 7.5 pp, medium confidence
86Gemini 3.5 Flash-Lite74.2%estimated ± 7.5 pp, medium confidence
87GLM-4.774.0%estimated ± 7.5 pp, medium confidence
88GLM-5V-Turbo74.0%estimated ± 7.5 pp, medium confidence
89Grok 4.2073.9%estimated ± 13.9 pp, low confidence
90Qwen 3.6 Max (preview)73.9%measured
91Ling 3.1 Flash73.8%estimated ± 7.5 pp, medium confidence
92DeepSeek V3.1 (Reasoning)73.7%estimated ± 7.5 pp, medium confidence
93Qwen3.7 Max73.6%measured
94GLM-5-Turbo73.2%estimated ± 7.5 pp, medium confidence
95GPT-4.172.7%estimated ± 7.5 pp, medium confidence
96Kimi K272.3%estimated ± 7.5 pp, medium confidence
97MiMo-V2.6-Flash72.0%estimated ± 7.5 pp, medium confidence
98Muse Glimmer 30B72.0%estimated ± 7.5 pp, medium confidence
99MiniMax M2.771.8%estimated ± 7.5 pp, medium confidence
100MiMo-V2-Pro71.6%estimated ± 7.5 pp, medium confidence
101Qwen3.6 Plus71.6%measured
102Qwen3.7 Plus71.4%measured
103Gemini 2.5 Flash71.1%estimated ± 7.5 pp, medium confidence
104Claude 4.1 Opus Thinking71.0%estimated ± 10.4 pp, low confidence
105Mistral Large 470.8%estimated ± 7.5 pp, medium confidence
106Step 3.7 Flash70.8%estimated ± 7.5 pp, medium confidence
107Seed 2.1 Pro70.8%measured
108GPT-5.4 nano70.7%estimated ± 7.5 pp, medium confidence
109Claude Opus 4.570.6%measured
110Solar Open 270.6%estimated ± 12.7 pp, low confidence
111DeepSeek V370.5%estimated ± 7.5 pp, medium confidence
112Qwen3.5 397B70.4%measured
113Grok 4.1 Fast (Reasoning)70.1%estimated ± 7.5 pp, medium confidence
114Mistral Large 370.0%estimated ± 7.5 pp, medium confidence
115Llama 4 Maverick69.9%estimated ± 7.5 pp, medium confidence
116Mistral Medium 3.5 128B69.7%estimated ± 7.5 pp, medium confidence
117Qwen3.5 397B (Reasoning)69.5%estimated ± 7.5 pp, medium confidence
118Qwen3.8-Flash-Next69.5%estimated ± 7.5 pp, medium confidence
119o3-pro69.5%estimated ± 10.4 pp, low confidence
120Qwen3 Max69.4%estimated ± 7.5 pp, medium confidence
121DeepSeek V3 032469.3%estimated ± 7.5 pp, medium confidence
122GLM-5.269.3%estimated ± 7.5 pp, medium confidence
123Nemotron 3 Super 100B69.3%estimated ± 7.5 pp, medium confidence
124Kimi K2.569.2%measured
125DeepSeek V3.269.0%estimated ± 7.5 pp, medium confidence
126GLM-5.168.7%estimated ± 7.5 pp, medium confidence
127Grok Code Fast 168.4%estimated ± 7.5 pp, medium confidence
128Llama 3.1 405B68.1%estimated ± 7.5 pp, medium confidence
129DeepSeek V3.168.0%estimated ± 7.5 pp, medium confidence
130Grok 4 Fast (Reasoning)67.6%estimated ± 7.5 pp, medium confidence
131Claude 4 Sonnet67.5%estimated ± 7.5 pp, medium confidence
132Seed 2.1 Turbo67.4%measured
133Ornith-1.5-35B-A3B67.3%estimated ± 9.1 pp, medium confidence
134Trinity-Large-Preview67.3%estimated ± 7.5 pp, medium confidence
135Trinity-Large-Thinking67.3%estimated ± 7.5 pp, medium confidence
136MiMo-V2.5-Pro67.2%estimated ± 7.5 pp, medium confidence
137Qwen3.5-122B-A10B67.1%measured
138GLM-566.8%measured
139Mercury 2.566.7%estimated ± 7.5 pp, medium confidence
140MAI-Thinking-166.5%estimated ± 9.2 pp, low confidence
141GPT-OSS 120B66.5%estimated ± 7.5 pp, medium confidence
142Mistral Small 466.3%estimated ± 7.5 pp, medium confidence
143Mistral Small 4 (Reasoning)66.3%estimated ± 7.5 pp, medium confidence
144Nemotron 3 Ultra66.2%estimated ± 7.5 pp, medium confidence
145Qwen3.6-27B66.0%measured
146GLM-4.666.0%estimated ± 7.5 pp, medium confidence
147Qwen3.5-27B65.6%measured
148Claude Sonnet 4.565.3%estimated ± 9.2 pp, low confidence
149Qwen3.6-35B-A3B64.7%measured
150GPT-4.1 mini64.6%estimated ± 7.5 pp, medium confidence
151Nemotron Ultra 253B64.3%estimated ± 7.5 pp, medium confidence
152Gemma 4 31B64.2%estimated ± 7.5 pp, medium confidence
153GPT-4o64.0%estimated ± 7.5 pp, medium confidence
154Mistral Large 264.0%estimated ± 7.5 pp, medium confidence
155Claude 4.1 Opus64.0%estimated ± 10.4 pp, low confidence
156Qwen3.5-35B-A3B63.4%measured
157MiMo-V2-Omni63.2%estimated ± 7.5 pp, medium confidence
158Gemma 4 26B A4B62.9%estimated ± 7.5 pp, medium confidence
159Ultravox v0.6 Llama 3.3 70B62.8%estimated ± 7.5 pp, medium confidence
160North Mini Code62.6%estimated ± 7.5 pp, medium confidence
161Solar Pro 462.6%estimated ± 7.5 pp, medium confidence
162Qwen3 235B 250762.6%measured
163A.X K262.2%estimated ± 7.5 pp, medium confidence
164Solar Pro 362.1%estimated ± 7.5 pp, medium confidence
165Mistral Medium 361.8%estimated ± 7.5 pp, medium confidence
166Ling 3.0 Flash61.6%estimated ± 7.5 pp, medium confidence
167Ling 3.0 Flash FP861.6%estimated ± 7.5 pp, medium confidence
168Ornith-1.5-9B61.1%estimated ± 9.1 pp, medium confidence
169Claude 3 Haiku60.7%estimated ± 7.5 pp, medium confidence
170Sarvam 105B60.7%estimated ± 7.5 pp, medium confidence
171Nemotron 3 Nano 30B60.2%estimated ± 7.5 pp, medium confidence
172Grok 4.1 Fast60.1%estimated ± 7.5 pp, medium confidence
173Nova Pro59.6%estimated ± 7.5 pp, medium confidence
174MiniMax M359.3%estimated ± 7.5 pp, medium confidence
175K-Exaone58.8%estimated ± 7.5 pp, medium confidence
176GLM-4.5-Air58.6%estimated ± 7.5 pp, medium confidence
177o1-pro58.5%estimated ± 9.2 pp, low confidence
178Solar Pro 258.3%estimated ± 7.5 pp, medium confidence
179Ternary Bonsai 2 27B58.2%estimated ± 4.5 pp, high confidence
180GPT-OSS 20B58.1%estimated ± 7.5 pp, medium confidence
181Gemma 4 12B57.4%estimated ± 7.5 pp, medium confidence
182Ling 2.6 Flash57.4%estimated ± 7.5 pp, medium confidence
183MiMo-V2-Flash57.4%estimated ± 7.5 pp, medium confidence
184Qwen3.8-27B57.4%estimated ± 7.5 pp, medium confidence
185Quasar 438B57.2%estimated ± 7.5 pp, medium confidence
186Llama 4 Scout56.7%estimated ± 7.5 pp, medium confidence
187Nemotron 3 Nano Omni 30B A3B56.7%estimated ± 7.5 pp, medium confidence
188LongCat-Flash-Lite-Sparse56.1%estimated ± 1.2 pp, low confidence
189o3-mini55.8%estimated ± 9.2 pp, low confidence
190Qwen3-Omni-30B-A3B-Thinking55.6%estimated ± 7.5 pp, medium confidence
191Ling 3.0 Flash VL55.2%estimated ± 7.5 pp, medium confidence
192Nemotron 3.5 Lightning 30B A3B NVFP455.2%estimated ± 7.5 pp, medium confidence
193Qwen3-Omni-30B-A3B-Instruct55.0%estimated ± 7.5 pp, medium confidence
194Phi-454.6%estimated ± 7.5 pp, medium confidence
195GPT-4.1 nano53.8%estimated ± 7.5 pp, medium confidence
196K-EXAONE 2.052.6%estimated ± 7.5 pp, medium confidence
197Gemma 3 27B52.4%estimated ± 7.5 pp, medium confidence
198Mellum2-12B-A2.5B-Thinking52.3%estimated ± 4.5 pp, high confidence
199Sarvam 30B51.5%estimated ± 7.5 pp, medium confidence
200Granite 4.2 8B48.3%estimated ± 7.5 pp, medium confidence
201o1-preview48.1%estimated ± 10.4 pp, low confidence
202Celeris-147.8%estimated ± 7.5 pp, medium confidence
203ZAYA1-8B47.4%estimated ± 9.2 pp, low confidence
204Exaone 4.0 32B46.8%estimated ± 7.5 pp, medium confidence
205Granite 4.2 30B45.5%estimated ± 7.5 pp, medium confidence
206LFM2.5-8B-A1B43.6%estimated ± 7.5 pp, medium confidence
207Granite 4.2 3B43.0%estimated ± 7.5 pp, medium confidence
208Command A+42.2%estimated ± 7.5 pp, medium confidence
209Gemma 4 E4B41.3%estimated ± 7.5 pp, medium confidence
210Ling 3.0 Tiny41.0%estimated ± 7.5 pp, medium confidence
211MiniCPM5-2B40.8%measured
212Claude 3 Opus40.2%estimated ± 10.4 pp, low confidence
213DeepSeek R1 Distill Qwen 32B39.1%estimated ± 10.4 pp, low confidence
214Gemini 1.5 Pro37.5%estimated ± 10.4 pp, low confidence
215Claude 3.5 Sonnet36.4%estimated ± 9.2 pp, low confidence
216Mellum2-12B-A2.5B-Instruct35.9%estimated ± 4.5 pp, high confidence
217ZAYA1-74B-Preview35.1%estimated ± 9.2 pp, low confidence
218Gemma 4 E2B34.7%estimated ± 7.5 pp, low confidence
219GPT-4 Turbo34.4%estimated ± 10.4 pp, low confidence
220Qwen2.5 Coder 32B Instruct33.3%estimated ± 10.4 pp, low confidence
221GPT-4o mini33.0%estimated ± 10.4 pp, low confidence
222LFM2.5-VL-1.6B-Extract31.7%estimated ± 7.5 pp, low confidence
223Soofi S 30B-A3B30.4%estimated ± 9.2 pp, low confidence
224Phi-4 Multimodal Instruct29.8%estimated ± 10.4 pp, low confidence
225Granite-4.0-H-1B29.3%estimated ± 7.5 pp, low confidence
226LFM2.5-VL-450M29.3%estimated ± 9.2 pp, low confidence
227LFM2.5-230M29.3%estimated ± 9.2 pp, low confidence
228Exaone 4.0 1.2B28.4%estimated ± 7.5 pp, low confidence
229Gemini 1.0 Pro27.9%estimated ± 10.4 pp, low confidence
230LFM2.5-2.6B25.8%estimated ± 7.5 pp, low confidence
231LLaDA2.2-mini23.7%estimated ± 13.9 pp, low confidence
232Granite-4.0-350M23.5%estimated ± 7.5 pp, low confidence
233MiniCPM5-1B23.1%measured
234Granite-4.0-H-350M23.0%estimated ± 7.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General