benchgap
Agentic · tools

Toolathlon-Verified leaderboard

As of 2026-10-07, the highest measured score on Toolathlon-Verified is 80.6% by Claude Opus 5. 143 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 580.6%measured
2GPT-6 Astra80.0%estimated ± 1.0 pp, low confidence
3Gemini 4 Argon79.8%estimated ± 1.0 pp, medium confidence
4Muse Spark 1.379.7%estimated ± 1.0 pp, medium confidence
5GPT-5.6 Sol79.5%estimated ± 1.0 pp, medium confidence
6GPT-6 Sol79.4%estimated ± 1.0 pp, medium confidence
7Gemini 3.8 Flash79.3%estimated ± 1.0 pp, medium confidence
8GPT-5.6 Terra78.7%estimated ± 1.0 pp, medium confidence
9Gemini 3.7 Flash78.5%estimated ± 1.0 pp, medium confidence
10GLM-5.3-Flash78.4%measured
11GPT-5.6 Luna78.3%estimated ± 1.0 pp, medium confidence
12Claude Fable 5.177.8%measured
13Claude Opus 5.577.8%measured
14Claude Sonnet 5.577.8%measured
15GPT-6.1 Sol76.9%estimated ± 1.9 pp, medium confidence
16MiMo-V2.6-Pro76.9%measured
17GPT-6 Luna76.9%estimated ± 2.2 pp, medium confidence
18Claude Fable 576.7%estimated ± 1.9 pp, medium confidence
19Claude Haiku 5.576.6%estimated ± 2.2 pp, medium confidence
20Grok 4.776.4%estimated ± 2.2 pp, medium confidence
21Mistral Large 476.1%estimated ± 2.2 pp, medium confidence
22Grok 4.675.9%estimated ± 2.3 pp, medium confidence
23GPT-5.5 Pro74.5%estimated ± 7.7 pp, medium confidence
24DeepSeek V4.1 Flash74.4%estimated ± 2.2 pp, medium confidence
25Ling 3.1 Flash74.3%estimated ± 2.2 pp, low confidence
26DeepSeek V4 Pro 081374.1%measured
27Hy4 preview74.1%measured
28Step 5 Preview74.1%measured
29Fugu Cyber74.1%estimated ± 2.2 pp, low confidence
30Gemini 3.8 Flash Cyber74.0%estimated ± 2.2 pp, low confidence
31Qwen3.8 Max Preview74.0%estimated ± 2.3 pp, medium confidence
32GPT-5.4 Pro73.7%estimated ± 7.7 pp, medium confidence
33MiMo-V2.6-Flash73.6%measured
34Claude Mythos 573.5%estimated ± 2.2 pp, low confidence
35Qwen3.8-Flash-Next73.5%measured
36Gemini 3.5 Flash Cyber73.4%estimated ± 2.2 pp, low confidence
37Claude Opus 4.873.4%estimated ± 1.0 pp, medium confidence
38Claude Mythos Preview73.4%estimated ± 2.2 pp, low confidence
39Kimi K373.2%measured
40Atria Dawn Preview73.2%estimated ± 1.3 pp, low confidence
41Claude 4.1 Opus73.2%estimated ± 1.3 pp, low confidence
42Claude 4 Sonnet73.2%estimated ± 1.3 pp, low confidence
43Claude Haiku 4.573.2%estimated ± 1.3 pp, low confidence
44Claude Opus 4.573.2%estimated ± 1.3 pp, low confidence
45Claude Opus 4.673.2%estimated ± 1.3 pp, low confidence
46Claude Sonnet 4.573.2%estimated ± 1.3 pp, low confidence
47Gemini 3 Flash73.2%estimated ± 1.3 pp, low confidence
48Gemini 3 Pro73.2%estimated ± 1.3 pp, low confidence
49GPT-5.1-Codex73.2%estimated ± 1.3 pp, low confidence
50GPT-5.273.2%estimated ± 1.3 pp, low confidence
51GPT-5.2-Codex73.2%estimated ± 1.3 pp, low confidence
52GPT-5.3 Codex73.2%estimated ± 1.3 pp, low confidence
53GPT-5.473.2%estimated ± 1.3 pp, low confidence
54GPT-5 (high)73.2%estimated ± 1.3 pp, low confidence
55Kimi K2.573.2%estimated ± 1.3 pp, low confidence
56Qwen3.5 Plus73.2%estimated ± 1.3 pp, low confidence
57Qwen3.8-27B73.2%estimated ± 1.3 pp, low confidence
58GLM-5.373.0%measured
59Muse Glimmer 30B73.0%estimated ± 2.2 pp, low confidence
60Qwen3.8 Max72.5%measured
61Claude Opus 4.7 (Adaptive)72.3%estimated ± 1.0 pp, low confidence
62Ornith-1.5-397B71.2%measured
63Inkling71.0%estimated ± 1.9 pp, low confidence
64Claude Sonnet 570.9%estimated ± 2.3 pp, medium confidence
65Muse Spark 1.270.7%estimated ± 2.3 pp, medium confidence
66DeepSeek V4 Flash 073170.3%measured
67GLM-5.170.2%estimated ± 2.2 pp, low confidence
68Muse Spark 1.169.8%estimated ± 1.0 pp, low confidence
69Claude Opus 4.769.6%estimated ± 1.0 pp, low confidence
70Grok 4.569.5%estimated ± 2.3 pp, medium confidence
71GPT-5.568.8%estimated ± 1.0 pp, low confidence
72GLM-5.267.6%estimated ± 2.3 pp, medium confidence
73Nemotron 3 Ultra67.1%estimated ± 1.9 pp, low confidence
74Agents-A163.9%estimated ± 5.9 pp, medium confidence
75Claude Sonnet 4.662.8%estimated ± 1.0 pp, low confidence
76Quasar 438B62.3%estimated ± 2.3 pp, medium confidence
77Beam61.9%estimated ± 7.7 pp, medium confidence
78Muse Spark61.4%estimated ± 2.2 pp, low confidence
79GLM-561.2%estimated ± 2.2 pp, low confidence
80Gemini 3.6 Flash60.0%estimated ± 2.3 pp, medium confidence
81Apodex 1.159.6%estimated ± 2.4 pp, high confidence
82Apodex 1.1 Mini59.6%estimated ± 2.4 pp, high confidence
83Ling 3.0 Flash VL58.1%estimated ± 2.4 pp, high confidence
84Gemini 3.5 Flash57.0%estimated ± 2.3 pp, medium confidence
85Qwen3.5 397B57.0%estimated ± 5.5 pp, low confidence
86dots3-note Preview55.6%measured
87Solar Pro 455.5%estimated ± 2.4 pp, medium confidence
88Hy355.3%estimated ± 2.3 pp, medium confidence
89Hy3 Preview55.3%estimated ± 2.3 pp, medium confidence
90Inkling-Small54.4%measured
91Qwen3.7 Max53.2%estimated ± 2.3 pp, low confidence
92Kimi K2.652.5%estimated ± 1.0 pp, low confidence
93MiniMax M352.5%estimated ± 1.0 pp, low confidence
94MiMo-V2.5-Pro51.7%estimated ± 2.3 pp, low confidence
95Kimi K2.7 Code51.5%estimated ± 2.3 pp, low confidence
96Agents-A1-4B51.4%estimated ± 7.7 pp, medium confidence
97GLM-4.750.4%estimated ± 2.4 pp, medium confidence
98Step 3.7 Flash50.4%estimated ± 2.4 pp, medium confidence
99Laguna S 2.149.7%measured
100Ling 3.0 Flash49.5%estimated ± 2.3 pp, low confidence
101Ling 3.0 Flash FP849.5%estimated ± 2.3 pp, low confidence
102Qwen3.6 Plus49.1%estimated ± 2.4 pp, medium confidence
103Ornith-1.5-35B-A3B48.7%measured
104Qwen3.6-27B48.3%estimated ± 2.3 pp, low confidence
105GPT-5.4 mini47.7%estimated ± 2.3 pp, low confidence
106A.X K247.3%estimated ± 2.4 pp, medium confidence
107Qwen3.5-27B45.7%estimated ± 7.7 pp, medium confidence
108Qwen3.5-35B-A3B45.7%estimated ± 7.7 pp, medium confidence
109GPT-5.4 nano44.8%estimated ± 2.3 pp, low confidence
110Grok 4.344.0%estimated ± 2.3 pp, low confidence
111MiniMax M2.743.3%estimated ± 2.3 pp, low confidence
112Solar Open 242.8%estimated ± 9.3 pp, medium confidence
113Qwen3.7 Plus42.5%estimated ± 1.0 pp, low confidence
114Gemini 3.5 Flash-Lite41.9%estimated ± 2.3 pp, low confidence
115Ornith-1.5-9B41.2%measured
116Qwen3.6-35B-A3B40.3%estimated ± 2.3 pp, low confidence
117Kimi K2.5 (Reasoning)38.9%estimated ± 2.4 pp, medium confidence
118GPT-5.137.8%estimated ± 2.4 pp, medium confidence
119LongCat-Flash-Lite-Sparse33.4%estimated ± 7.7 pp, low confidence
120LLaDA2.2-flash31.4%estimated ± 9.3 pp, low confidence
121Gemini 3.1 Pro31.1%estimated ± 2.3 pp, low confidence
122Qwen3.5-122B-A10B29.5%estimated ± 2.3 pp, low confidence
123Mistral Medium 3.5 128B28.9%estimated ± 2.3 pp, low confidence
124MiniCPM5-2B27.0%estimated ± 2.4 pp, medium confidence
125Gemma 4 31B22.4%estimated ± 2.3 pp, low confidence
126GPT-OSS 120B20.9%estimated ± 2.3 pp, low confidence
127Nemotron 3.5 Lightning 30B A3B NVFP420.8%estimated ± 2.3 pp, low confidence
128MiMo-V2-Flash18.5%estimated ± 2.4 pp, medium confidence
129Nemotron 3 Super 100B14.8%estimated ± 2.3 pp, low confidence
130Granite 4.2 8B13.5%estimated ± 2.3 pp, low confidence
131Command A+13.2%estimated ± 2.3 pp, low confidence
132Gemini 2.5 Pro13.0%estimated ± 2.3 pp, low confidence
133Granite 4.2 30B11.6%estimated ± 2.4 pp, medium confidence
134Gemma 4 26B A4B10.3%estimated ± 2.4 pp, medium confidence
135Ling 3.0 Tiny9.7%estimated ± 2.4 pp, medium confidence
136Mistral Large 39.2%estimated ± 2.3 pp, low confidence
137Mistral Small 45.6%estimated ± 2.3 pp, low confidence
138Mistral Small 4 (Reasoning)5.6%estimated ± 2.3 pp, low confidence
139GPT-OSS 20B5.5%estimated ± 2.3 pp, low confidence
140Trinity-Large-Preview4.6%estimated ± 2.3 pp, low confidence
141Trinity-Large-Thinking4.6%estimated ± 2.3 pp, low confidence
142Nemotron 3 Nano 30B4.0%estimated ± 2.3 pp, low confidence
143DeepSeek V33.2%estimated ± 2.3 pp, low confidence
144Celeris-12.6%estimated ± 2.3 pp, low confidence
145Llama 4 Maverick2.5%estimated ± 2.3 pp, low confidence
146Llama 4 Scout2.2%estimated ± 2.3 pp, low confidence
147Gemma 3 27B0.6%estimated ± 2.3 pp, low confidence
148DeepSeek V3 03240.0%estimated ± 2.4 pp, medium confidence
149Gemma 4 12B0.0%estimated ± 2.4 pp, medium confidence
150Gemma 4 E2B0.0%estimated ± 2.4 pp, medium confidence
151Gemma 4 E4B0.0%estimated ± 2.4 pp, medium confidence
152GPT-4.1 mini0.0%estimated ± 2.4 pp, medium confidence
153GPT-4.1 nano0.0%estimated ± 2.4 pp, medium confidence
154GPT-4o0.0%estimated ± 2.4 pp, medium confidence
155GPT-4o mini0.0%estimated ± 2.4 pp, medium confidence
156Granite 4.2 3B0.0%estimated ± 2.4 pp, medium confidence
157K-Exaone0.0%estimated ± 2.4 pp, medium confidence
158LFM2.5-2.6B0.0%estimated ± 2.4 pp, medium confidence
159Ling 2.6 Flash0.0%estimated ± 2.4 pp, medium confidence
160Mercury 2.50.0%estimated ± 2.4 pp, medium confidence
161Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 2.4 pp, medium confidence
162North Mini Code0.0%estimated ± 2.4 pp, medium confidence
163Solar Pro 30.0%estimated ± 2.4 pp, medium confidence
164Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 2.4 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General