benchgap
Knowledge & reasoning

HLE leaderboard

As of 2026-10-07, the highest measured score on HLE is 65.0% by Claude Fable 5.1. 185 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.567.6%estimated ± 1.9 pp, high confidence
2Claude Fable 5.165.0%measured
3Claude Opus 564.7%measured
4Claude Mythos 564.5%measured
5Claude Sonnet 5.563.7%estimated ± 5.7 pp, medium confidence
6Muse Spark 1.162.1%measured
7Claude Fable 561.7%estimated ± 6.3 pp, medium confidence
8GPT-6 Astra60.9%estimated ± 5.5 pp, low confidence
9GPT-6.1 Sol60.9%estimated ± 5.5 pp, low confidence
10Gemini 3.8 Flash60.7%estimated ± 5.5 pp, low confidence
11GPT-5.6 Sol60.7%estimated ± 5.5 pp, low confidence
12GPT-6 Sol60.5%estimated ± 5.5 pp, low confidence
13Grok 4.660.0%estimated ± 5.5 pp, low confidence
14GPT-5.4 Pro58.7%measured
15GPT-5.6 Terra58.6%estimated ± 5.5 pp, low confidence
16Sakana Fugu-Ultra58.3%estimated ± 5.7 pp, medium confidence
17Claude Opus 4.857.9%measured
18Gemini 3.7 Flash57.8%estimated ± 6.3 pp, medium confidence
19Pareto 26.957.4%estimated ± 5.7 pp, medium confidence
20Claude Sonnet 557.4%measured
21GPT-5.5 Pro57.2%measured
22Gemini 3.1 Pro56.6%estimated ± 5.5 pp, low confidence
23Apodex 1.156.1%measured
24Kimi K356.0%measured
25Sakana Fugu55.9%estimated ± 5.7 pp, medium confidence
26Gemini 4 Argon55.5%estimated ± 9.5 pp, medium confidence
27Hy4 preview55.4%measured
28MiMo-V2.6-Pro55.2%estimated ± 9.5 pp, medium confidence
29Grok 4.555.0%estimated ± 5.5 pp, low confidence
30Claude Haiku 5.554.8%estimated ± 5.7 pp, medium confidence
31Claude Opus 4.7 (Adaptive)54.7%measured
32GLM-5.254.7%measured
33Muse Spark 1.354.4%estimated ± 9.5 pp, medium confidence
34Claude Opus 4.653.0%measured
35dots3-note Preview52.6%measured
36GLM-5.152.3%measured
37GPT-5.552.2%measured
38GPT-5.452.1%measured
39GPT-5.6 Luna51.6%estimated ± 5.5 pp, low confidence
40GLM-5.351.1%estimated ± 9.5 pp, medium confidence
41GLM-550.4%measured
42Muse Spark50.4%measured
43Ling 3.1 Flash50.4%estimated ± 9.5 pp, medium confidence
44Grok 4.750.2%estimated ± 9.5 pp, medium confidence
45Muse Spark 1.250.2%estimated ± 9.5 pp, medium confidence
46Qwen3.8 Max Preview50.2%estimated ± 9.5 pp, medium confidence
47GPT-5.3 Codex49.6%estimated ± 9.5 pp, medium confidence
48Claude Sonnet 4.649.0%measured
49GLM-5.3-Flash48.6%estimated ± 9.5 pp, medium confidence
50Claude Opus 4.6 (Adaptive)48.5%estimated ± 6.3 pp, medium confidence
51MiMo-V2.5-Pro48.0%measured
52Inkling-Small47.8%measured
53Agents-A147.6%measured
54Pareto 26.10 Preview47.0%estimated ± 9.2 pp, medium confidence
55Step 5 Preview46.5%measured
56Gemini 3.6 Flash46.0%estimated ± 6.3 pp, medium confidence
57Inkling46.0%measured
58GPT-6 Luna46.0%estimated ± 5.5 pp, low confidence
59Beam45.9%estimated ± 5.7 pp, medium confidence
60MiMo-V2.6-Flash45.8%estimated ± 9.5 pp, medium confidence
61GPT-5.244.8%estimated ± 6.3 pp, medium confidence
62Grok 4.2044.8%estimated ± 5.5 pp, low confidence
63Ornith-1.5-397B44.6%measured
64Mistral Large 444.4%estimated ± 9.5 pp, medium confidence
65Gemini 3 Pro Deep Think44.2%estimated ± 6.3 pp, medium confidence
66Claude Opus 4.5 Thinking44.0%estimated ± 6.3 pp, medium confidence
67Gemini 3 Pro43.9%estimated ± 6.3 pp, medium confidence
68Claude Sonnet 4.543.8%estimated ± 6.3 pp, medium confidence
69Claude Sonnet 4.5 Thinking43.8%estimated ± 6.3 pp, medium confidence
70Kimi K2.7 Code43.8%estimated ± 9.5 pp, medium confidence
71Qwen3.8 Max43.6%measured
72Gemini 3.1 Flash-Lite43.5%estimated ± 9.8 pp, low confidence
73Quasar 438B43.2%estimated ± 9.5 pp, medium confidence
74DeepSeek V4 Pro 081342.7%measured
75GPT-5.2-Codex42.4%estimated ± 9.5 pp, medium confidence
76GPT-5.142.2%estimated ± 6.4 pp, medium confidence
77GPT-5.4 mini41.5%measured
78Qwen3.7 Max41.4%measured
79Gemini 3.5 Flash40.2%measured
80Qwen3.5 Flash39.5%estimated ± 9.8 pp, low confidence
81GPT-5.1-Codex38.5%estimated ± 9.5 pp, medium confidence
82GPT-5.1-Codex-Max38.5%estimated ± 9.5 pp, medium confidence
83GPT-5 (high)38.5%estimated ± 9.5 pp, medium confidence
84Solar Open 238.3%estimated ± 5.7 pp, medium confidence
85Interfaze Beta37.9%estimated ± 8.2 pp, medium confidence
86Solar Pro 437.8%estimated ± 9.2 pp, medium confidence
87GPT-5.4 nano37.7%measured
88Claude Opus 4.737.5%estimated ± 9.5 pp, medium confidence
89MiMo-V2.537.3%estimated ± 9.8 pp, low confidence
90Hy337.2%estimated ± 9.5 pp, medium confidence
91DeepSeek V4.1 Flash36.8%measured
92Apodex 1.1 Mini36.7%estimated ± 9.5 pp, medium confidence
93Qwen3.8-Omni-Flash36.5%measured
94Qwen3.8-Flash-Next35.9%measured
95Qwen 3.6 Max (preview)35.5%estimated ± 7.8 pp, medium confidence
96Grok 4.335.0%measured
97MiniMax M335.0%estimated ± 9.5 pp, medium confidence
98DeepSeek V4 Flash 073134.8%measured
99Kimi K2.634.7%measured
100Qwen3.7 Plus34.7%measured
101GLM-4.534.1%estimated ± 9.8 pp, low confidence
102Nemotron 3 Super 100B33.7%estimated ± 9.5 pp, medium confidence
103MiniMax M2.733.3%estimated ± 9.2 pp, medium confidence
104Grok 4.1 Fast (Reasoning)33.2%estimated ± 9.5 pp, medium confidence
105Grok 4 Fast (Reasoning)33.2%estimated ± 9.5 pp, medium confidence
106Seed 2.1 Pro33.1%estimated ± 7.8 pp, medium confidence
107Muse Glimmer 30B32.5%estimated ± 9.5 pp, medium confidence
108Step 3.7 Flash31.7%estimated ± 9.5 pp, medium confidence
109Kimi K2.5 (Reasoning)31.7%estimated ± 8.2 pp, medium confidence
110DeepSeek V3.1 (Reasoning)30.9%estimated ± 9.5 pp, medium confidence
111Grok 430.9%estimated ± 9.5 pp, medium confidence
112Ling 3.0 Flash VL30.9%estimated ± 9.5 pp, medium confidence
113Claude Opus 4.530.8%measured
114Qwen3.8-27B30.8%measured
115Seed 2.1 Turbo30.5%estimated ± 7.8 pp, medium confidence
116Qwen3.5-122B-A10B30.2%estimated ± 7.8 pp, medium confidence
117Kimi K2.530.1%measured
118Claude Haiku 4.529.8%estimated ± 9.8 pp, low confidence
119Qwen3.5-27B29.1%estimated ± 7.8 pp, medium confidence
120DeepSeek-R129.0%estimated ± 9.5 pp, medium confidence
121Gemini 2.5 Flash29.0%estimated ± 9.5 pp, medium confidence
122Gemini 3 Flash29.0%estimated ± 9.5 pp, medium confidence
123GPT-OSS 20B29.0%estimated ± 9.5 pp, medium confidence
124Qwen3.6 Plus28.8%measured
125Qwen3.5 397B28.7%measured
126Claude 4 Sonnet27.8%estimated ± 9.5 pp, medium confidence
127GPT-OSS 120B27.8%estimated ± 9.5 pp, medium confidence
128K-Exaone27.8%estimated ± 9.5 pp, medium confidence
129MiMo-V2-Omni27.8%estimated ± 9.5 pp, medium confidence
130o327.8%estimated ± 9.5 pp, medium confidence
131A.X K227.8%measured
132o3-pro27.5%estimated ± 9.9 pp, medium confidence
133Qwen3.5-35B-A3B27.4%estimated ± 7.8 pp, medium confidence
134DeepSeek V3.227.0%estimated ± 9.5 pp, medium confidence
135Nemotron 3 Nano 30B27.0%estimated ± 9.5 pp, medium confidence
136Qwen3.5 397B (Reasoning)27.0%estimated ± 9.5 pp, medium confidence
137Qwen3 235B 250726.8%estimated ± 7.8 pp, medium confidence
138Nemotron 3 Ultra26.7%measured
139Gemma 4 31B26.5%measured
140Ornith-1.5-35B-A3B25.6%measured
141Hy3 Preview25.5%measured
142GLM-5V-Turbo25.5%estimated ± 9.5 pp, medium confidence
143MAI-Thinking-125.1%estimated ± 8.2 pp, medium confidence
144Ling 3.0 Flash FP824.8%estimated ± 8.2 pp, medium confidence
145GLM-4.724.8%measured
146MiMo-V2-Flash24.4%estimated ± 8.2 pp, medium confidence
147Qwen3.6-27B24.0%measured
148Command A+23.5%estimated ± 9.5 pp, medium confidence
149GLM-5-Turbo23.5%estimated ± 9.5 pp, medium confidence
150MiMo-V2-Pro23.5%estimated ± 9.5 pp, medium confidence
151Mistral Small 423.5%estimated ± 9.5 pp, medium confidence
152Mistral Small 4 (Reasoning)23.5%estimated ± 9.5 pp, medium confidence
153North Mini Code23.5%estimated ± 9.5 pp, medium confidence
154Sarvam 30B23.5%estimated ± 9.5 pp, medium confidence
155Ling 3.0 Flash22.7%measured
156Qwen3.6-35B-A3B21.4%measured
157Claude 4.1 Opus20.3%estimated ± 9.9 pp, medium confidence
158Ornith-1.5-9B20.2%measured
159o1-preview19.4%estimated ± 9.9 pp, low confidence
160Claude 3 Opus19.4%estimated ± 9.9 pp, low confidence
161DeepSeek R1 Distill Qwen 32B19.4%estimated ± 9.9 pp, low confidence
162Gemini 1.5 Pro19.4%estimated ± 9.9 pp, low confidence
163GPT-4 Turbo19.4%estimated ± 9.9 pp, low confidence
164GPT-4o mini19.4%estimated ± 9.9 pp, low confidence
165Phi-4 Multimodal Instruct19.4%estimated ± 9.9 pp, low confidence
166Qwen2.5 Coder 32B Instruct19.4%estimated ± 9.9 pp, low confidence
167Gemini 1.0 Pro19.4%estimated ± 9.9 pp, low confidence
168Mercury 2.519.2%estimated ± 9.2 pp, medium confidence
169Celeris-119.0%estimated ± 9.5 pp, medium confidence
170Claude 3 Haiku19.0%estimated ± 9.5 pp, medium confidence
171Claude 4.1 Opus Thinking19.0%estimated ± 9.5 pp, medium confidence
172DeepSeek V3 032419.0%estimated ± 9.5 pp, medium confidence
173DeepSeek V3.119.0%estimated ± 9.5 pp, medium confidence
174Exaone 4.0 1.2B19.0%estimated ± 9.5 pp, medium confidence
175Exaone 4.0 32B19.0%estimated ± 9.5 pp, medium confidence
176Gemini 3.5 Flash-Lite19.0%estimated ± 9.5 pp, medium confidence
177Gemma 3 27B19.0%estimated ± 9.5 pp, medium confidence
178GLM-4.5-Air19.0%estimated ± 9.5 pp, medium confidence
179GLM-4.619.0%estimated ± 9.5 pp, medium confidence
180GPT-4o19.0%estimated ± 9.5 pp, medium confidence
181GPT-5 (medium)19.0%estimated ± 9.5 pp, medium confidence
182Granite-4.0-350M19.0%estimated ± 9.5 pp, medium confidence
183Granite-4.0-H-1B19.0%estimated ± 9.5 pp, medium confidence
184Granite-4.0-H-350M19.0%estimated ± 9.5 pp, medium confidence
185Grok 4.1 Fast19.0%estimated ± 9.5 pp, medium confidence
186Grok Code Fast 119.0%estimated ± 9.5 pp, medium confidence
187Kimi K219.0%estimated ± 9.5 pp, medium confidence
188LFM2.5-2.6B19.0%estimated ± 9.5 pp, medium confidence
189LFM2.5-8B-A1B19.0%estimated ± 9.5 pp, medium confidence
190LFM2.5-VL-1.6B-Extract19.0%estimated ± 9.5 pp, medium confidence
191Ling 3.0 Tiny19.0%estimated ± 9.5 pp, medium confidence
192Llama 3.1 405B19.0%estimated ± 9.5 pp, medium confidence
193Llama 4 Maverick19.0%estimated ± 9.5 pp, medium confidence
194Llama 4 Scout19.0%estimated ± 9.5 pp, medium confidence
195Mistral Large 219.0%estimated ± 9.5 pp, medium confidence
196Mistral Large 319.0%estimated ± 9.5 pp, medium confidence
197Mistral Medium 319.0%estimated ± 9.5 pp, medium confidence
198Mistral Medium 3.5 128B19.0%estimated ± 9.5 pp, medium confidence
199Nemotron Ultra 253B19.0%estimated ± 9.5 pp, medium confidence
200Nova Pro19.0%estimated ± 9.5 pp, medium confidence
201Phi-419.0%estimated ± 9.5 pp, medium confidence
202Qwen3 Max19.0%estimated ± 9.5 pp, medium confidence
203Qwen3-Omni-30B-A3B-Instruct19.0%estimated ± 9.5 pp, medium confidence
204Qwen3-Omni-30B-A3B-Thinking19.0%estimated ± 9.5 pp, medium confidence
205Sarvam 105B19.0%estimated ± 9.5 pp, medium confidence
206Solar Pro 219.0%estimated ± 9.5 pp, medium confidence
207Solar Pro 319.0%estimated ± 9.5 pp, medium confidence
208Ultravox v0.6 Llama 3.3 70B19.0%estimated ± 9.5 pp, medium confidence
209Gemini 2.5 Pro18.8%measured
210Ternary Bonsai 2 27B18.7%estimated ± 5.9 pp, low confidence
211o1-pro18.6%estimated ± 8.2 pp, low confidence
212K-EXAONE 2.018.3%measured
213Gemma 4 26B A4B17.2%measured
214o3-mini17.0%estimated ± 8.2 pp, low confidence
215Nemotron 3.5 Lightning 30B A3B NVFP415.8%estimated ± 5.7 pp, medium confidence
216o115.8%estimated ± 8.2 pp, low confidence
217Laguna XS.215.7%estimated ± 9.8 pp, low confidence
218Trinity-Large-Thinking15.6%estimated ± 9.2 pp, medium confidence
219Laguna M.115.4%estimated ± 9.8 pp, low confidence
220Nemotron 3 Nano Omni 30B A3B13.4%estimated ± 8.2 pp, low confidence
221ZAYA1-8B12.7%estimated ± 8.2 pp, low confidence
222Mellum2-12B-A2.5B-Thinking12.3%estimated ± 5.9 pp, low confidence
223LongCat-Flash-Lite-Sparse11.1%estimated ± 2.5 pp, low confidence
224Granite 4.2 30B10.4%estimated ± 8.2 pp, low confidence
225GPT-4.110.3%estimated ± 8.2 pp, low confidence
226GPT-4.1 mini9.5%estimated ± 8.2 pp, low confidence
227Granite 4.2 8B9.5%estimated ± 8.2 pp, low confidence
228MiniCPM5-2B8.9%measured
229Gemma 4 12B8.2%estimated ± 5.7 pp, low confidence
230Claude 3.5 Sonnet7.8%estimated ± 8.2 pp, low confidence
231DeepSeek V37.7%estimated ± 8.2 pp, low confidence
232Ling 2.6 Flash7.7%estimated ± 8.2 pp, low confidence
233Gemma 4 E4B7.6%estimated ± 8.2 pp, low confidence
234ZAYA1-74B-Preview7.2%estimated ± 8.2 pp, low confidence
235Granite 4.2 3B6.5%estimated ± 8.2 pp, low confidence
236GPT-4.1 nano5.5%estimated ± 8.2 pp, low confidence
237Gemma 4 E2B4.2%estimated ± 8.2 pp, low confidence
238Soofi S 30B-A3B4.2%estimated ± 8.2 pp, low confidence
239Trinity-Large-Preview2.8%estimated ± 9.2 pp, low confidence
240LFM2.5-VL-450M1.9%estimated ± 8.2 pp, low confidence
241LFM2.5-230M1.9%estimated ± 8.2 pp, low confidence
242LLaDA2.2-mini0.0%estimated ± 9.2 pp, low confidence
243Mellum2-12B-A2.5B-Instruct0.0%estimated ± 5.9 pp, low confidence
244MiniCPM5-1B0.0%estimated ± 5.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General