benchgap
Agentic · tools

HLE w/ tools leaderboard

As of 2026-10-07, the highest measured score on HLE w/ tools is 67.7% by Claude Opus 5.5. 131 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Muse Spark 1.373.5%estimated ± 4.6 pp, low confidence
2MiMo-V2.6-Flash72.7%estimated ± 5.0 pp, low confidence
3MiMo-V2.6-Pro71.7%estimated ± 5.0 pp, low confidence
4Claude Opus 5.567.7%measured
5Claude Fable 565.8%estimated ± 5.7 pp, low confidence
6Ling 3.1 Flash65.4%estimated ± 5.0 pp, low confidence
7Grok 4.765.2%estimated ± 4.6 pp, medium confidence
8Claude Opus 564.7%measured
9GLM-5.264.6%estimated ± 4.6 pp, medium confidence
10Claude Sonnet 5.564.5%measured
11Fugu Cyber64.2%estimated ± 5.0 pp, low confidence
12DeepSeek V4.1 Flash63.9%measured
13Gemini 3.8 Flash Cyber63.3%estimated ± 5.0 pp, low confidence
14GLM-5.362.5%measured
15Qwen3.8 Max Preview61.7%estimated ± 6.1 pp, medium confidence
16Gemini 4 Argon60.5%estimated ± 6.1 pp, medium confidence
17Grok 4.660.4%estimated ± 6.1 pp, medium confidence
18Gemini 3.7 Flash60.2%estimated ± 5.7 pp, low confidence
19DeepSeek V4 Pro 081360.0%measured
20GPT-5.6 Sol59.3%estimated ± 4.6 pp, medium confidence
21Gemini 3.5 Flash Cyber59.3%estimated ± 5.0 pp, low confidence
22Kimi K359.3%estimated ± 4.6 pp, high confidence
23GPT-6.1 Sol59.2%estimated ± 6.1 pp, medium confidence
24GPT-5.5 Pro59.1%estimated ± 4.6 pp, high confidence
25Claude Mythos Preview59.1%estimated ± 5.0 pp, low confidence
26GPT-5.4 Pro59.0%estimated ± 4.6 pp, high confidence
27Step 5 Preview58.9%estimated ± 4.6 pp, high confidence
28Claude Mythos 558.8%estimated ± 4.6 pp, high confidence
29GPT-5.6 Terra58.7%estimated ± 4.6 pp, high confidence
30GPT-6 Sol58.3%estimated ± 4.6 pp, medium confidence
31Claude Fable 5.158.2%estimated ± 4.6 pp, medium confidence
32Qwen3.8-Flash-Next58.0%estimated ± 5.9 pp, medium confidence
33GPT-5.557.6%estimated ± 4.6 pp, high confidence
34Claude Opus 4.857.5%estimated ± 4.6 pp, high confidence
35Claude Haiku 5.557.4%measured
36Claude Sonnet 557.4%measured
37Claude Opus 4.657.2%estimated ± 4.6 pp, high confidence
38GPT-6 Astra57.2%measured
39MiniMax M357.1%estimated ± 4.6 pp, high confidence
40GPT-5.6 Luna57.0%estimated ± 4.6 pp, high confidence
41Muse Spark 1.256.7%estimated ± 6.1 pp, medium confidence
42GPT-5.456.6%estimated ± 4.6 pp, high confidence
43Qwen3.8 Max56.2%measured
44Apodex 1.156.1%measured
45Ornith-1.5-397B56.1%measured
46GPT-6 Luna55.7%estimated ± 6.1 pp, medium confidence
47Atria Dawn Preview55.6%estimated ± 2.4 pp, medium confidence
48Gemini 3.8 Flash55.4%estimated ± 4.6 pp, low confidence
49Hy4 preview55.4%measured
50Mistral Large 455.3%estimated ± 6.1 pp, medium confidence
51Qwen3.8-27B55.3%estimated ± 6.1 pp, medium confidence
52GLM-5.3-Flash55.3%measured
53Kimi K2.655.0%estimated ± 2.4 pp, medium confidence
54Grok 4.554.4%estimated ± 6.1 pp, medium confidence
55Qwen3.7 Max53.5%measured
56Gemini 3.5 Flash53.2%estimated ± 6.1 pp, medium confidence
57Claude Opus 4.7 (Adaptive)53.1%estimated ± 4.6 pp, high confidence
58dots3-note Preview52.6%measured
59Gemini 3.6 Flash51.8%estimated ± 6.1 pp, medium confidence
60Inkling-Small50.3%estimated ± 4.6 pp, high confidence
61Beam50.3%estimated ± 4.6 pp, high confidence
62Hy3 Preview50.0%estimated ± 6.1 pp, medium confidence
63Inkling49.8%estimated ± 4.6 pp, high confidence
64Claude Opus 4.549.6%estimated ± 2.4 pp, medium confidence
65Apodex 1.1 Mini49.4%estimated ± 6.1 pp, medium confidence
66Quasar 438B49.1%estimated ± 6.1 pp, medium confidence
67Ling 3.0 Flash VL48.6%estimated ± 6.1 pp, medium confidence
68MiMo-V2.5-Pro47.6%estimated ± 6.1 pp, medium confidence
69Agents-A147.6%measured
70Step 3.7 Flash47.2%measured
71Grok 4.346.6%estimated ± 6.1 pp, medium confidence
72Hy346.1%estimated ± 6.1 pp, medium confidence
73Kimi K2.7 Code45.4%estimated ± 6.1 pp, medium confidence
74DeepSeek V4 Flash 073145.1%measured
75Qwen3.6 Plus45.0%estimated ± 2.4 pp, medium confidence
76GPT-5.4 mini44.8%estimated ± 6.1 pp, medium confidence
77MiniMax M2.744.8%estimated ± 6.1 pp, low confidence
78Qwen3.5 397B44.3%estimated ± 2.4 pp, medium confidence
79Qwen3.6-27B44.1%estimated ± 6.1 pp, low confidence
80Gemini 3.5 Flash-Lite44.1%estimated ± 6.1 pp, low confidence
81A.X K243.5%estimated ± 6.1 pp, low confidence
82Ling 3.0 Flash43.4%estimated ± 2.4 pp, medium confidence
83GPT-5.4 nano43.1%estimated ± 6.1 pp, low confidence
84Ling 3.0 Flash FP843.1%estimated ± 6.1 pp, low confidence
85GPT-5 (high)42.4%estimated ± 6.1 pp, low confidence
86Kimi K2.541.3%estimated ± 2.4 pp, medium confidence
87GPT-5.140.1%estimated ± 6.1 pp, low confidence
88Gemini 3.1 Pro39.1%estimated ± 6.1 pp, low confidence
89Muse Glimmer 30B39.0%estimated ± 6.1 pp, low confidence
90Qwen3.7 Plus38.5%estimated ± 6.1 pp, low confidence
91Mistral Medium 3.5 128B38.4%estimated ± 6.1 pp, low confidence
92Laguna S 2.138.3%estimated ± 5.9 pp, medium confidence
93Nemotron 3 Ultra37.4%measured
94MiniCPM5-2B37.0%estimated ± 6.1 pp, low confidence
95GLM-5.136.6%estimated ± 4.6 pp, high confidence
96Agents-A1-4B35.7%estimated ± 4.6 pp, high confidence
97GLM-535.7%estimated ± 2.4 pp, medium confidence
98GPT-5.235.2%estimated ± 4.6 pp, high confidence
99MiMo-V2-Flash35.0%estimated ± 6.1 pp, low confidence
100Gemma 4 31B34.7%estimated ± 6.1 pp, low confidence
101GPT-OSS 120B34.5%estimated ± 6.1 pp, low confidence
102Qwen3.5-122B-A10B34.4%estimated ± 4.6 pp, high confidence
103Qwen3.5-27B33.8%estimated ± 4.6 pp, high confidence
104Qwen3.5-35B-A3B33.8%estimated ± 4.6 pp, high confidence
105Kimi K2.5 (Reasoning)33.7%estimated ± 4.6 pp, high confidence
106Granite 4.2 30B33.6%estimated ± 6.1 pp, low confidence
107Solar Open 233.5%estimated ± 7.2 pp, medium confidence
108Ornith-1.5-35B-A3B33.4%measured
109Gemma 4 26B A4B33.3%estimated ± 6.1 pp, low confidence
110GLM-4.733.2%estimated ± 4.6 pp, high confidence
111Ling 3.0 Tiny33.2%estimated ± 6.1 pp, low confidence
112Solar Pro 433.2%estimated ± 4.6 pp, high confidence
113LongCat-Flash-Lite-Sparse33.2%estimated ± 4.6 pp, high confidence
114Nemotron 3.5 Lightning 30B A3B NVFP433.2%estimated ± 4.6 pp, medium confidence
115Command A+32.0%estimated ± 6.1 pp, low confidence
116Celeris-131.6%estimated ± 6.1 pp, low confidence
117DeepSeek V331.6%estimated ± 6.1 pp, low confidence
118DeepSeek V3 032431.6%estimated ± 6.1 pp, low confidence
119Gemini 2.5 Pro31.6%estimated ± 6.1 pp, low confidence
120Gemma 3 27B31.6%estimated ± 6.1 pp, low confidence
121Gemma 4 12B31.6%estimated ± 6.1 pp, low confidence
122Gemma 4 E2B31.6%estimated ± 6.1 pp, low confidence
123Gemma 4 E4B31.6%estimated ± 6.1 pp, low confidence
124GPT-4.1 mini31.6%estimated ± 6.1 pp, low confidence
125GPT-4.1 nano31.6%estimated ± 6.1 pp, low confidence
126GPT-4o31.6%estimated ± 6.1 pp, low confidence
127GPT-4o mini31.6%estimated ± 6.1 pp, low confidence
128GPT-OSS 20B31.6%estimated ± 6.1 pp, low confidence
129Granite 4.2 3B31.6%estimated ± 6.1 pp, low confidence
130Granite 4.2 8B31.6%estimated ± 6.1 pp, low confidence
131K-Exaone31.6%estimated ± 6.1 pp, low confidence
132LFM2.5-2.6B31.6%estimated ± 6.1 pp, low confidence
133Ling 2.6 Flash31.6%estimated ± 6.1 pp, low confidence
134Llama 4 Maverick31.6%estimated ± 6.1 pp, low confidence
135Llama 4 Scout31.6%estimated ± 6.1 pp, low confidence
136Mercury 2.531.6%estimated ± 6.1 pp, low confidence
137Mistral Large 331.6%estimated ± 6.1 pp, low confidence
138Mistral Small 431.6%estimated ± 6.1 pp, low confidence
139Mistral Small 4 (Reasoning)31.6%estimated ± 6.1 pp, low confidence
140Nemotron 3 Nano 30B31.6%estimated ± 6.1 pp, low confidence
141Nemotron 3 Nano Omni 30B A3B31.6%estimated ± 6.1 pp, low confidence
142Nemotron 3 Super 100B31.6%estimated ± 6.1 pp, low confidence
143North Mini Code31.6%estimated ± 6.1 pp, low confidence
144Solar Pro 331.6%estimated ± 6.1 pp, low confidence
145Trinity-Large-Preview31.6%estimated ± 6.1 pp, low confidence
146Trinity-Large-Thinking31.6%estimated ± 6.1 pp, low confidence
147Ultravox v0.6 Llama 3.3 70B31.6%estimated ± 6.1 pp, low confidence
148Qwen3.6-35B-A3B30.6%estimated ± 2.4 pp, medium confidence
149Ornith-1.5-9B30.5%measured
150Claude Sonnet 4.628.1%estimated ± 5.0 pp, low confidence
151LLaDA2.2-flash23.9%estimated ± 7.2 pp, low confidence
152Muse Spark 1.118.0%estimated ± 5.0 pp, low confidence
153Muse Spark3.5%estimated ± 5.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General