benchgap
Knowledge & reasoning

HLE w/o tools leaderboard

As of 2026-10-07, the highest measured score on HLE w/o tools is 64.4% by Claude Opus 5.5. 200 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.564.4%measured
2Gemini 4 Argon61.3%estimated ± 5.7 pp, medium confidence
3Claude Fable 5.160.9%measured
4Claude Mythos 559.0%measured
5Claude Sonnet 5.556.9%measured
6Claude Opus 556.3%measured
7Claude Fable 554.9%estimated ± 5.1 pp, medium confidence
8GPT-6.1 Sol52.7%estimated ± 5.0 pp, low confidence
9Gemini 3.8 Flash52.5%estimated ± 5.0 pp, medium confidence
10Muse Spark 1.152.2%measured
11GPT-6 Sol52.1%estimated ± 5.0 pp, medium confidence
12Gemini 3.7 Flash51.9%estimated ± 5.1 pp, medium confidence
13Grok 4.651.4%estimated ± 5.0 pp, medium confidence
14Claude Opus 4.751.3%estimated ± 5.1 pp, medium confidence
15MiMo-V2.6-Pro50.4%estimated ± 5.7 pp, medium confidence
16Sakana Fugu-Ultra50.0%measured
17Claude Opus 4.849.8%measured
18Muse Spark 1.349.4%estimated ± 5.7 pp, medium confidence
19Gemini 3.6 Flash49.1%estimated ± 5.1 pp, medium confidence
20Pareto 26.949.0%measured
21GPT-6 Astra48.5%estimated ± 4.4 pp, medium confidence
22Sakana Fugu47.2%measured
23Claude Opus 4.7 (Adaptive)46.9%measured
24GPT-5.6 Sol46.6%estimated ± 4.4 pp, high confidence
25Gemini 3 Flash46.1%estimated ± 5.1 pp, medium confidence
26Claude Haiku 5.545.9%measured
27Apodex 1.145.7%estimated ± 5.3 pp, medium confidence
28Gemini 3.1 Pro45.4%measured
29Step 5 Preview44.7%estimated ± 4.4 pp, high confidence
30Muse Spark 1.244.6%estimated ± 5.1 pp, medium confidence
31Ornith-1.5-397B44.6%measured
32Grok 4.544.5%estimated ± 5.0 pp, medium confidence
33Qwen3.8 Max43.6%measured
34Kimi K343.5%measured
35GPT-5.6 Terra43.5%estimated ± 4.4 pp, high confidence
36Hy4 preview43.4%measured
37Claude Sonnet 543.2%measured
38GPT-5.5 Pro43.1%measured
39dots3-note Preview42.9%estimated ± 5.3 pp, medium confidence
40Muse Spark42.8%measured
41GPT-5.4 Pro42.7%measured
42Pareto 26.10 Preview42.4%estimated ± 4.4 pp, high confidence
43Grok 4.742.2%estimated ± 5.7 pp, medium confidence
44Qwen3.8 Max Preview42.2%estimated ± 5.7 pp, medium confidence
45GPT-5.6 Luna42.1%estimated ± 4.4 pp, high confidence
46GPT-5.241.9%estimated ± 4.9 pp, high confidence
47GPT-5.3 Codex41.4%estimated ± 5.7 pp, medium confidence
48GPT-5.541.4%measured
49GLM-5.240.5%measured
50Claude Opus 4.640.0%measured
51GPT-5.439.8%measured
52Qwen3.7 Max39.6%estimated ± 2.9 pp, high confidence
53Gemini 3.5 Flash39.0%estimated ± 2.9 pp, high confidence
54Agents-A138.8%estimated ± 5.3 pp, medium confidence
55Claude Opus 4.6 (Adaptive)38.3%estimated ± 5.7 pp, medium confidence
56Gemini 3 Pro38.0%estimated ± 5.7 pp, medium confidence
57Ling 3.1 Flash37.7%estimated ± 5.7 pp, medium confidence
58GLM-5.337.0%estimated ± 5.1 pp, medium confidence
59Gemini 3 Pro Deep Think36.3%estimated ± 6.9 pp, medium confidence
60Beam36.2%measured
61Qwen3.8-Flash-Next35.9%measured
62Interfaze Beta35.6%estimated ± 4.4 pp, high confidence
63Qwen3.6 Plus35.4%estimated ± 4.9 pp, high confidence
64Qwen3.8-Omni-Flash35.2%estimated ± 2.9 pp, high confidence
65DeepSeek V4.1 Flash35.1%estimated ± 2.9 pp, high confidence
66DeepSeek V4 Pro 081334.6%estimated ± 2.9 pp, high confidence
67Grok 4.334.5%estimated ± 4.9 pp, high confidence
68Gemini 3.1 Flash-Lite34.2%estimated ± 5.1 pp, medium confidence
69MiMo-V2.5-Pro34.0%measured
70Claude Sonnet 4.633.9%estimated ± 4.9 pp, high confidence
71GLM-5.3-Flash33.8%estimated ± 5.1 pp, medium confidence
72GPT-6 Luna33.8%estimated ± 5.0 pp, medium confidence
73Kimi K2.633.7%estimated ± 2.9 pp, high confidence
74GPT-5.2-Codex33.4%estimated ± 5.7 pp, medium confidence
75Qwen3.7 Plus33.2%estimated ± 2.9 pp, high confidence
76Solar Pro 432.8%estimated ± 4.4 pp, high confidence
77MiMo-V2.6-Flash32.7%estimated ± 5.7 pp, medium confidence
78Gemini 3.5 Flash-Lite32.6%estimated ± 5.1 pp, medium confidence
79Kimi K2.7 Code32.6%estimated ± 5.7 pp, medium confidence
80Mistral Large 432.6%estimated ± 5.7 pp, medium confidence
81Grok 4.2031.6%measured
82Inkling-Small31.6%measured
83Apodex 1.1 Mini31.6%estimated ± 5.7 pp, medium confidence
84Hy330.9%estimated ± 5.7 pp, medium confidence
85Qwen3.8-27B30.8%measured
86Inkling30.0%measured
87Qwen3.5 397B29.3%estimated ± 4.9 pp, high confidence
88Solar Open 228.8%measured
89MiniMax M328.3%estimated ± 5.1 pp, medium confidence
90Claude Sonnet 4.5 Thinking28.2%estimated ± 6.0 pp, medium confidence
91GPT-5.4 mini28.2%measured
92DeepSeek V4 Flash 073128.2%estimated ± 2.9 pp, high confidence
93Qwen3.5 Flash28.1%estimated ± 5.1 pp, medium confidence
94Qwen 3.6 Max (preview)27.9%estimated ± 5.7 pp, medium confidence
95Qwen3.6-27B27.6%estimated ± 4.9 pp, high confidence
96MiMo-V2-Pro27.5%estimated ± 5.7 pp, medium confidence
97o3-pro27.5%estimated ± 6.3 pp, medium confidence
98Claude Opus 4.5 Thinking27.2%estimated ± 5.7 pp, medium confidence
99Kimi K2.5 (Reasoning)27.0%estimated ± 4.9 pp, high confidence
100GLM-527.0%estimated ± 2.9 pp, high confidence
101Nemotron 3 Ultra26.7%measured
102MiMo-V2.526.5%estimated ± 5.1 pp, medium confidence
103MiniMax M2.726.3%estimated ± 4.4 pp, high confidence
104GLM-4.626.0%estimated ± 5.1 pp, medium confidence
105Kimi K2.525.9%estimated ± 2.9 pp, high confidence
106Ornith-1.5-35B-A3B25.6%measured
107GLM-4.525.6%estimated ± 5.1 pp, medium confidence
108GPT-5.125.5%estimated ± 5.7 pp, medium confidence
109GPT-5 (high)25.5%estimated ± 5.7 pp, medium confidence
110Claude Opus 4.525.4%estimated ± 4.9 pp, high confidence
111Claude Haiku 4.525.2%estimated ± 5.1 pp, medium confidence
112Mistral Medium 3.5 128B25.1%estimated ± 5.1 pp, low confidence
113Laguna M.125.1%estimated ± 5.1 pp, low confidence
114Laguna XS.225.1%estimated ± 5.1 pp, low confidence
115GLM-5-Turbo24.8%estimated ± 5.7 pp, medium confidence
116Qwen3.5-122B-A10B24.3%estimated ± 4.9 pp, high confidence
117GPT-5.4 nano24.3%measured
118Hy3 Preview24.0%estimated ± 2.9 pp, high confidence
119Claude 4.1 Opus23.8%estimated ± 6.3 pp, medium confidence
120GLM-5.123.8%estimated ± 4.4 pp, high confidence
121Grok 423.6%estimated ± 5.7 pp, medium confidence
122Qwen3.6-35B-A3B22.9%estimated ± 4.9 pp, high confidence
123GPT-5.1-Codex22.6%estimated ± 5.7 pp, medium confidence
124GPT-5.1-Codex-Max22.6%estimated ± 5.7 pp, medium confidence
125Ternary Bonsai 2 27B22.5%estimated ± 4.4 pp, high confidence
126GPT-5 (medium)22.3%estimated ± 5.7 pp, medium confidence
127GLM-4.722.1%estimated ± 4.9 pp, high confidence
128A.X K222.0%estimated ± 4.4 pp, high confidence
129Qwen3.5-27B21.7%estimated ± 4.9 pp, high confidence
130Ornith-1.5-9B20.2%measured
131Gemma 4 31B19.5%measured
132MiMo-V2-Omni19.1%estimated ± 5.7 pp, medium confidence
133Ling 3.0 Flash VL19.0%estimated ± 5.7 pp, medium confidence
134Muse Glimmer 30B19.0%estimated ± 5.7 pp, medium confidence
135Qwen3.5-35B-A3B18.9%estimated ± 4.9 pp, high confidence
136Step 3.7 Flash18.4%estimated ± 5.7 pp, medium confidence
137Ling 3.0 Flash18.2%estimated ± 2.9 pp, high confidence
138MAI-Thinking-118.2%estimated ± 4.4 pp, high confidence
139MiMo-V2-Flash17.9%estimated ± 4.9 pp, high confidence
140Nemotron 3 Super 100B17.8%estimated ± 5.7 pp, medium confidence
141Ling 3.0 Flash FP817.7%estimated ± 4.4 pp, high confidence
142Claude Sonnet 4.517.4%estimated ± 4.9 pp, high confidence
143o317.2%estimated ± 5.7 pp, medium confidence
144Qwen3.5 397B (Reasoning)16.9%estimated ± 5.7 pp, medium confidence
145GPT-OSS 120B16.7%estimated ± 5.7 pp, medium confidence
146Gemini 2.5 Pro16.7%estimated ± 4.9 pp, high confidence
147Grok 4.1 Fast (Reasoning)16.4%estimated ± 5.7 pp, medium confidence
148Grok 4 Fast (Reasoning)16.2%estimated ± 5.7 pp, medium confidence
149Quasar 438B15.8%estimated ± 5.7 pp, medium confidence
150o1-preview15.3%estimated ± 6.3 pp, low confidence
151GLM-5V-Turbo14.4%estimated ± 5.7 pp, medium confidence
152K-EXAONE 2.013.8%estimated ± 4.4 pp, high confidence
153DeepSeek-R113.2%estimated ± 5.7 pp, medium confidence
154DeepSeek V3.1 (Reasoning)11.8%estimated ± 5.7 pp, medium confidence
155K-Exaone11.5%estimated ± 5.7 pp, medium confidence
156o1-pro11.4%estimated ± 4.9 pp, high confidence
157Nemotron 3.5 Lightning 30B A3B NVFP410.5%measured
158Claude 4.1 Opus Thinking10.2%estimated ± 5.7 pp, medium confidence
159Qwen3 235B 250710.2%estimated ± 4.9 pp, high confidence
160o3-mini9.9%estimated ± 4.9 pp, high confidence
161Command A+9.8%estimated ± 5.7 pp, medium confidence
162Qwen3 Max9.7%estimated ± 5.7 pp, medium confidence
163Mercury 2.59.5%estimated ± 4.4 pp, high confidence
164Nemotron 3 Nano 30B9.3%estimated ± 5.7 pp, medium confidence
165DeepSeek V3.29.1%estimated ± 5.7 pp, medium confidence
166North Mini Code9.0%estimated ± 5.7 pp, medium confidence
167o19.0%estimated ± 4.9 pp, high confidence
168GPT-OSS 20B8.9%estimated ± 5.7 pp, medium confidence
169Sarvam 105B8.9%estimated ± 5.7 pp, medium confidence
170Gemma 4 26B A4B8.7%measured
171Solar Pro 38.3%estimated ± 5.7 pp, low confidence
172Mistral Small 48.0%estimated ± 5.7 pp, low confidence
173Mistral Small 4 (Reasoning)8.0%estimated ± 5.7 pp, low confidence
174Trinity-Large-Thinking7.6%estimated ± 4.4 pp, high confidence
175Ling 3.0 Tiny7.5%estimated ± 5.7 pp, low confidence
176Granite 4.2 30B6.6%estimated ± 4.9 pp, medium confidence
177GPT-4.16.5%estimated ± 4.9 pp, medium confidence
178Grok Code Fast 16.4%estimated ± 5.7 pp, low confidence
179GPT-4.1 mini6.4%estimated ± 4.9 pp, medium confidence
180Granite 4.2 8B6.4%estimated ± 4.9 pp, medium confidence
181Nemotron 3 Nano Omni 30B A3B6.4%estimated ± 4.4 pp, medium confidence
182ZAYA1-8B6.2%estimated ± 4.4 pp, medium confidence
183Claude 3.5 Sonnet6.2%estimated ± 4.9 pp, medium confidence
184DeepSeek V36.2%estimated ± 4.9 pp, medium confidence
185Ling 2.6 Flash6.2%estimated ± 4.9 pp, medium confidence
186Gemma 4 E4B6.1%estimated ± 4.9 pp, medium confidence
187MiniCPM5-2B6.1%estimated ± 4.4 pp, medium confidence
188Granite 4.2 3B6.1%estimated ± 4.9 pp, medium confidence
189GPT-4.1 nano6.0%estimated ± 4.9 pp, medium confidence
190LongCat-Flash-Lite-Sparse6.0%estimated ± 4.4 pp, medium confidence
191Gemma 4 E2B6.0%estimated ± 4.9 pp, medium confidence
192LFM2.5-VL-450M6.0%estimated ± 4.9 pp, medium confidence
193Qwen3-Omni-30B-A3B-Thinking6.0%estimated ± 5.7 pp, low confidence
194Sarvam 30B6.0%estimated ± 5.7 pp, low confidence
195Kimi K25.9%estimated ± 5.7 pp, low confidence
196Nemotron Ultra 253B5.9%estimated ± 5.7 pp, low confidence
197Trinity-Large-Preview5.8%estimated ± 4.4 pp, medium confidence
198Mellum2-12B-A2.5B-Thinking5.8%estimated ± 4.4 pp, medium confidence
199ZAYA1-74B-Preview5.7%estimated ± 4.4 pp, medium confidence
200LLaDA2.2-mini5.7%estimated ± 4.4 pp, medium confidence
201Soofi S 30B-A3B5.7%estimated ± 4.4 pp, medium confidence
202Mellum2-12B-A2.5B-Instruct5.7%estimated ± 4.4 pp, medium confidence
203LFM2.5-230M5.7%estimated ± 4.4 pp, medium confidence
204MiniCPM5-1B5.7%estimated ± 4.4 pp, medium confidence
205GLM-4.5-Air5.6%estimated ± 5.7 pp, low confidence
206LFM2.5-8B-A1B5.5%estimated ± 5.7 pp, low confidence
207Celeris-15.4%estimated ± 5.7 pp, low confidence
208DeepSeek V3.15.3%estimated ± 5.7 pp, low confidence
209Gemma 4 12B5.2%measured
210Granite-4.0-H-350M5.1%estimated ± 5.7 pp, low confidence
211LFM2.5-2.6B4.9%estimated ± 5.7 pp, low confidence
212Exaone 4.0 1.2B4.5%estimated ± 5.7 pp, low confidence
213Granite-4.0-350M4.3%estimated ± 5.7 pp, low confidence
214Grok 4.1 Fast4.0%estimated ± 5.7 pp, low confidence
215LFM2.5-VL-1.6B-Extract4.0%estimated ± 5.7 pp, low confidence
216Exaone 4.0 32B3.9%estimated ± 5.7 pp, low confidence
217Granite-4.0-H-1B3.9%estimated ± 5.7 pp, low confidence
218Phi-4 Multimodal Instruct3.9%estimated ± 5.7 pp, low confidence
219Llama 4 Maverick3.9%estimated ± 5.7 pp, low confidence
220DeepSeek V3 03243.7%estimated ± 5.7 pp, low confidence
221Gemini 2.5 Flash3.7%estimated ± 5.7 pp, low confidence
222DeepSeek R1 Distill Qwen 32B3.6%estimated ± 5.7 pp, low confidence
223Gemini 1.5 Pro3.6%estimated ± 5.7 pp, low confidence
224Qwen3-Omni-30B-A3B-Instruct3.6%estimated ± 5.7 pp, low confidence
225Gemma 3 27B3.5%estimated ± 5.7 pp, low confidence
226Claude 4 Sonnet3.4%estimated ± 5.7 pp, low confidence
227Gemini 1.0 Pro3.3%estimated ± 5.7 pp, low confidence
228GPT-4o mini3.3%estimated ± 5.7 pp, low confidence
229Mistral Large 33.3%estimated ± 5.7 pp, low confidence
230Claude 3 Haiku3.2%estimated ± 5.7 pp, low confidence
231Mistral Medium 33.2%estimated ± 5.7 pp, low confidence
232Llama 3.1 405B3.1%estimated ± 5.7 pp, low confidence
233Llama 4 Scout3.0%estimated ± 5.7 pp, low confidence
234Phi-43.0%estimated ± 5.7 pp, low confidence
235Solar Pro 22.9%estimated ± 5.7 pp, low confidence
236Ultravox v0.6 Llama 3.3 70B2.8%estimated ± 5.7 pp, low confidence
237Qwen2.5 Coder 32B Instruct2.7%estimated ± 5.7 pp, low confidence
238Mistral Large 22.6%estimated ± 5.7 pp, low confidence
239Nova Pro2.5%estimated ± 5.7 pp, low confidence
240GPT-4 Turbo2.4%estimated ± 5.7 pp, low confidence
241Claude 3 Opus2.2%estimated ± 5.7 pp, low confidence
242GPT-4o1.9%estimated ± 5.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General