benchgap
Knowledge & reasoning

HLE-Verified leaderboard

As of 2026-10-07, the highest measured score on HLE-Verified is 54.9% by Gemini 3.8 Flash. 206 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 3.8 Flash54.9%measured
2Gemini 4 Argon54.8%estimated ± 1.2 pp, low confidence
3Claude Fable 5.154.6%estimated ± 0.7 pp, medium confidence
4Claude Opus 5.554.6%estimated ± 0.7 pp, medium confidence
5GPT-6.1 Sol54.6%estimated ± 0.7 pp, low confidence
6GPT-6 Astra54.6%estimated ± 0.7 pp, low confidence
7GPT-6 Sol54.6%estimated ± 0.7 pp, medium confidence
8Claude Fable 554.6%estimated ± 0.7 pp, medium confidence
9GPT-5.6 Sol54.5%measured
10Claude Opus 554.4%measured
11Qwen3.8 Max54.3%estimated ± 4.6 pp, medium confidence
12GPT-5.554.1%estimated ± 0.7 pp, medium confidence
13Muse Spark 1.354.1%estimated ± 2.0 pp, medium confidence
14Claude Haiku 5.554.0%estimated ± 6.8 pp, low confidence
15Gemini 3.7 Flash53.6%measured
16MiniMax M353.0%estimated ± 2.0 pp, medium confidence
17Qwen3.8 Max Preview52.7%estimated ± 2.0 pp, medium confidence
18GPT-5.5 Pro52.5%estimated ± 0.7 pp, medium confidence
19GPT-5.6 Terra51.1%measured
20Claude Sonnet 5.550.6%estimated ± 6.0 pp, low confidence
21Qwen3.7 Max49.9%estimated ± 2.0 pp, medium confidence
22Qwen3.8-Flash-Next49.9%estimated ± 2.0 pp, medium confidence
23GPT-5.4 Pro46.4%estimated ± 0.7 pp, low confidence
24Grok 4.742.5%estimated ± 6.0 pp, low confidence
25GLM-5.342.3%estimated ± 2.0 pp, medium confidence
26DeepSeek V4.1 Flash41.3%estimated ± 6.0 pp, low confidence
27GPT-5.3 Codex38.7%estimated ± 2.0 pp, medium confidence
28dots3-note Preview35.9%estimated ± 0.7 pp, low confidence
29Step 5 Preview35.8%estimated ± 6.0 pp, low confidence
30Gemini 3.1 Pro35.3%estimated ± 0.7 pp, low confidence
31Claude Opus 4.5 Thinking35.3%estimated ± 0.7 pp, low confidence
32Claude Opus 4.6 (Adaptive)35.3%estimated ± 0.7 pp, low confidence
33Claude Opus 4.7 (Adaptive)35.3%estimated ± 0.7 pp, low confidence
34Claude Opus 4.835.3%estimated ± 0.7 pp, low confidence
35Claude Sonnet 4.535.3%estimated ± 0.7 pp, low confidence
36Claude Sonnet 4.5 Thinking35.3%estimated ± 0.7 pp, low confidence
37Claude Sonnet 4.635.3%estimated ± 0.7 pp, low confidence
38DeepSeek V4 Flash 073135.3%estimated ± 0.7 pp, low confidence
39DeepSeek V4 Pro 081335.3%estimated ± 0.7 pp, low confidence
40Gemini 3.5 Flash35.3%estimated ± 0.7 pp, low confidence
41Gemini 3.6 Flash35.3%estimated ± 0.7 pp, low confidence
42Gemini 3 Pro35.3%estimated ± 0.7 pp, low confidence
43Gemini 3 Pro Deep Think35.3%estimated ± 0.7 pp, low confidence
44GPT-5.235.3%estimated ± 0.7 pp, low confidence
45GPT-5.435.3%estimated ± 0.7 pp, low confidence
46GPT-5.4 mini35.3%estimated ± 0.7 pp, low confidence
47GPT-5.4 nano35.3%estimated ± 0.7 pp, low confidence
48GPT-5.6 Luna35.3%estimated ± 0.7 pp, low confidence
49GPT-6 Luna35.3%estimated ± 0.7 pp, low confidence
50Grok 4.2035.3%estimated ± 0.7 pp, low confidence
51Grok 4.535.3%estimated ± 0.7 pp, low confidence
52Grok 4.635.3%estimated ± 0.7 pp, low confidence
53Inkling-Small35.3%estimated ± 0.7 pp, low confidence
54Kimi K335.3%estimated ± 0.7 pp, low confidence
55Muse Spark35.3%estimated ± 0.7 pp, low confidence
56GLM-5.3-Flash32.9%estimated ± 2.0 pp, medium confidence
57Claude Sonnet 531.0%measured
58Kimi K2.631.0%estimated ± 2.0 pp, medium confidence
59MiMo-V2.6-Pro28.8%estimated ± 6.0 pp, low confidence
60Ling 3.1 Flash23.3%estimated ± 6.0 pp, low confidence
61Qwen3.8-27B21.9%estimated ± 2.0 pp, low confidence
62MiMo-V2.6-Flash21.4%estimated ± 6.0 pp, low confidence
63Muse Spark 1.220.8%estimated ± 2.0 pp, low confidence
64Mistral Large 420.3%estimated ± 6.0 pp, low confidence
65Grok 4.318.6%estimated ± 2.0 pp, low confidence
66Qwen3.7 Plus18.0%estimated ± 2.0 pp, low confidence
67GPT-5.2-Codex17.6%estimated ± 2.0 pp, low confidence
68Muse Spark 1.117.2%estimated ± 2.0 pp, low confidence
69Mercury 2.516.9%estimated ± 6.0 pp, low confidence
70Hy316.8%estimated ± 2.0 pp, low confidence
71Hy3 Preview16.8%estimated ± 2.0 pp, low confidence
72Kimi K2.7 Code16.6%estimated ± 2.0 pp, low confidence
73GLM-5.216.4%estimated ± 2.0 pp, low confidence
74Solar Pro 415.8%estimated ± 2.0 pp, low confidence
75Qwen 3.6 Max (preview)15.6%estimated ± 2.0 pp, low confidence
76Claude Opus 4.715.5%estimated ± 2.0 pp, low confidence
77Qwen3.6 Plus15.4%estimated ± 2.0 pp, low confidence
78Kimi K2.515.4%estimated ± 2.0 pp, low confidence
79Kimi K2.5 (Reasoning)15.4%estimated ± 2.0 pp, low confidence
80Grok 415.3%estimated ± 2.0 pp, low confidence
81MiniMax M2.715.3%estimated ± 2.0 pp, low confidence
82GPT-5.115.3%estimated ± 2.0 pp, low confidence
83Inkling15.3%estimated ± 2.0 pp, low confidence
84MiMo-V2-Pro15.3%estimated ± 2.0 pp, low confidence
85GLM-5.115.3%estimated ± 2.0 pp, low confidence
86Nemotron 3 Ultra15.3%estimated ± 2.0 pp, low confidence
87MiMo-V2.5-Pro15.3%estimated ± 2.0 pp, low confidence
88Apodex 1.115.3%estimated ± 2.0 pp, low confidence
89Apodex 1.1 Mini15.3%estimated ± 2.0 pp, low confidence
90Ling 3.0 Flash VL15.3%estimated ± 2.0 pp, low confidence
91Qwen3.5 397B15.3%estimated ± 2.0 pp, low confidence
92Qwen3.5 397B (Reasoning)15.3%estimated ± 2.0 pp, low confidence
93GPT-5.1-Codex15.3%estimated ± 2.0 pp, low confidence
94GPT-5.1-Codex-Max15.3%estimated ± 2.0 pp, low confidence
95GLM-4.715.3%estimated ± 2.0 pp, low confidence
96Qwen3.5-27B15.3%estimated ± 2.0 pp, low confidence
97A.X K215.3%estimated ± 2.0 pp, low confidence
98Gemma 4 31B15.3%estimated ± 2.0 pp, low confidence
99Qwen3.5-122B-A10B15.3%estimated ± 2.0 pp, low confidence
100GPT-5 (high)15.3%estimated ± 2.0 pp, low confidence
101Ling 3.0 Flash15.3%estimated ± 2.0 pp, low confidence
102Ling 3.0 Flash FP815.3%estimated ± 2.0 pp, low confidence
103Grok 4.1 Fast (Reasoning)15.3%estimated ± 2.0 pp, low confidence
104Celeris-115.3%estimated ± 2.0 pp, low confidence
105Claude 3 Haiku15.3%estimated ± 2.0 pp, low confidence
106Claude 3 Opus15.3%estimated ± 2.0 pp, low confidence
107Claude 4.1 Opus Thinking15.3%estimated ± 2.0 pp, low confidence
108Claude 4 Sonnet15.3%estimated ± 2.0 pp, low confidence
109Claude Opus 4.515.3%estimated ± 2.0 pp, low confidence
110Claude Opus 4.615.3%estimated ± 2.0 pp, low confidence
111Command A+15.3%estimated ± 2.0 pp, low confidence
112DeepSeek-R115.3%estimated ± 2.0 pp, low confidence
113DeepSeek R1 Distill Qwen 32B15.3%estimated ± 2.0 pp, low confidence
114DeepSeek V315.3%estimated ± 2.0 pp, low confidence
115DeepSeek V3 032415.3%estimated ± 2.0 pp, low confidence
116DeepSeek V3.115.3%estimated ± 2.0 pp, low confidence
117DeepSeek V3.1 (Reasoning)15.3%estimated ± 2.0 pp, low confidence
118DeepSeek V3.215.3%estimated ± 2.0 pp, low confidence
119Exaone 4.0 1.2B15.3%estimated ± 2.0 pp, low confidence
120Exaone 4.0 32B15.3%estimated ± 2.0 pp, low confidence
121Gemini 1.0 Pro15.3%estimated ± 2.0 pp, low confidence
122Gemini 1.5 Pro15.3%estimated ± 2.0 pp, low confidence
123Gemini 2.5 Flash15.3%estimated ± 2.0 pp, low confidence
124Gemini 2.5 Pro15.3%estimated ± 2.0 pp, low confidence
125Gemini 3.5 Flash-Lite15.3%estimated ± 2.0 pp, low confidence
126Gemini 3 Flash15.3%estimated ± 2.0 pp, low confidence
127Gemma 3 27B15.3%estimated ± 2.0 pp, low confidence
128Gemma 4 12B15.3%estimated ± 2.0 pp, low confidence
129Gemma 4 26B A4B15.3%estimated ± 2.0 pp, low confidence
130Gemma 4 E2B15.3%estimated ± 2.0 pp, low confidence
131Gemma 4 E4B15.3%estimated ± 2.0 pp, low confidence
132GLM-4.5-Air15.3%estimated ± 2.0 pp, low confidence
133GLM-4.615.3%estimated ± 2.0 pp, low confidence
134GLM-515.3%estimated ± 2.0 pp, low confidence
135GLM-5-Turbo15.3%estimated ± 2.0 pp, low confidence
136GLM-5V-Turbo15.3%estimated ± 2.0 pp, low confidence
137GPT-4.115.3%estimated ± 2.0 pp, low confidence
138GPT-4.1 mini15.3%estimated ± 2.0 pp, low confidence
139GPT-4.1 nano15.3%estimated ± 2.0 pp, low confidence
140GPT-4o15.3%estimated ± 2.0 pp, low confidence
141GPT-4o mini15.3%estimated ± 2.0 pp, low confidence
142GPT-5 (medium)15.3%estimated ± 2.0 pp, low confidence
143GPT-OSS 120B15.3%estimated ± 2.0 pp, low confidence
144GPT-OSS 20B15.3%estimated ± 2.0 pp, low confidence
145Granite-4.0-350M15.3%estimated ± 2.0 pp, low confidence
146Granite-4.0-H-1B15.3%estimated ± 2.0 pp, low confidence
147Granite-4.0-H-350M15.3%estimated ± 2.0 pp, low confidence
148Granite 4.2 30B15.3%estimated ± 2.0 pp, low confidence
149Granite 4.2 3B15.3%estimated ± 2.0 pp, low confidence
150Granite 4.2 8B15.3%estimated ± 2.0 pp, low confidence
151Grok 4.1 Fast15.3%estimated ± 2.0 pp, low confidence
152Grok 4 Fast (Reasoning)15.3%estimated ± 2.0 pp, low confidence
153Grok Code Fast 115.3%estimated ± 2.0 pp, low confidence
154K-Exaone15.3%estimated ± 2.0 pp, low confidence
155K-EXAONE 2.015.3%estimated ± 2.0 pp, low confidence
156Kimi K215.3%estimated ± 2.0 pp, low confidence
157LFM2.5-2.6B15.3%estimated ± 2.0 pp, low confidence
158LFM2.5-8B-A1B15.3%estimated ± 2.0 pp, low confidence
159LFM2.5-VL-1.6B-Extract15.3%estimated ± 2.0 pp, low confidence
160Ling 2.6 Flash15.3%estimated ± 2.0 pp, low confidence
161Ling 3.0 Tiny15.3%estimated ± 2.0 pp, low confidence
162Llama 3.1 405B15.3%estimated ± 2.0 pp, low confidence
163Llama 4 Maverick15.3%estimated ± 2.0 pp, low confidence
164Llama 4 Scout15.3%estimated ± 2.0 pp, low confidence
165MiMo-V2-Flash15.3%estimated ± 2.0 pp, low confidence
166MiMo-V2-Omni15.3%estimated ± 2.0 pp, low confidence
167MiniCPM5-2B15.3%estimated ± 2.0 pp, low confidence
168Mistral Large 215.3%estimated ± 2.0 pp, low confidence
169Mistral Large 315.3%estimated ± 2.0 pp, low confidence
170Mistral Medium 315.3%estimated ± 2.0 pp, low confidence
171Mistral Medium 3.5 128B15.3%estimated ± 2.0 pp, low confidence
172Mistral Small 415.3%estimated ± 2.0 pp, low confidence
173Mistral Small 4 (Reasoning)15.3%estimated ± 2.0 pp, low confidence
174Muse Glimmer 30B15.3%estimated ± 2.0 pp, low confidence
175Nemotron 3.5 Lightning 30B A3B NVFP415.3%estimated ± 2.0 pp, low confidence
176Nemotron 3 Nano 30B15.3%estimated ± 2.0 pp, low confidence
177Nemotron 3 Nano Omni 30B A3B15.3%estimated ± 2.0 pp, low confidence
178Nemotron 3 Super 100B15.3%estimated ± 2.0 pp, low confidence
179Nemotron Ultra 253B15.3%estimated ± 2.0 pp, low confidence
180North Mini Code15.3%estimated ± 2.0 pp, low confidence
181Nova Pro15.3%estimated ± 2.0 pp, low confidence
182o115.3%estimated ± 2.0 pp, low confidence
183o1-preview15.3%estimated ± 2.0 pp, low confidence
184o315.3%estimated ± 2.0 pp, low confidence
185o3-mini15.3%estimated ± 2.0 pp, low confidence
186o3-pro15.3%estimated ± 2.0 pp, low confidence
187Phi-415.3%estimated ± 2.0 pp, low confidence
188Phi-4 Multimodal Instruct15.3%estimated ± 2.0 pp, low confidence
189Quasar 438B15.3%estimated ± 2.0 pp, low confidence
190Qwen2.5 Coder 32B Instruct15.3%estimated ± 2.0 pp, low confidence
191Qwen3.5-35B-A3B15.3%estimated ± 2.0 pp, low confidence
192Qwen3.6-27B15.3%estimated ± 2.0 pp, low confidence
193Qwen3.6-35B-A3B15.3%estimated ± 2.0 pp, low confidence
194Qwen3 Max15.3%estimated ± 2.0 pp, low confidence
195Qwen3-Omni-30B-A3B-Instruct15.3%estimated ± 2.0 pp, low confidence
196Qwen3-Omni-30B-A3B-Thinking15.3%estimated ± 2.0 pp, low confidence
197Sarvam 105B15.3%estimated ± 2.0 pp, low confidence
198Sarvam 30B15.3%estimated ± 2.0 pp, low confidence
199Solar Pro 215.3%estimated ± 2.0 pp, low confidence
200Solar Pro 315.3%estimated ± 2.0 pp, low confidence
201Step 3.7 Flash15.3%estimated ± 2.0 pp, low confidence
202Trinity-Large-Preview15.3%estimated ± 2.0 pp, low confidence
203Trinity-Large-Thinking15.3%estimated ± 2.0 pp, low confidence
204Ultravox v0.6 Llama 3.3 70B15.3%estimated ± 2.0 pp, low confidence
205Qwen3.5 Flash0.0%estimated ± 4.6 pp, low confidence
206MiMo-V2.50.0%estimated ± 4.6 pp, low confidence
207Gemini 3.1 Flash-Lite0.0%estimated ± 4.6 pp, low confidence
208Claude Haiku 4.50.0%estimated ± 4.6 pp, low confidence
209GLM-4.50.0%estimated ± 4.6 pp, low confidence
210GPT-4 Turbo0.0%estimated ± 6.8 pp, low confidence
211Laguna M.10.0%estimated ± 4.6 pp, low confidence
212Laguna XS.20.0%estimated ± 4.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General