benchgap
Agentic · tools

Toolathlon leaderboard

As of 2026-10-07, the highest measured score on Toolathlon is 75.6% by Muse Spark 1.1. 154 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Mercury 2.590.3%estimated ± 6.0 pp, low confidence
2Mistral Medium 3.5 128B83.3%estimated ± 6.0 pp, low confidence
3Muse Spark 1.175.6%measured
4Claude Fable 5.167.2%estimated ± 1.7 pp, low confidence
5GPT-6 Astra66.4%estimated ± 1.7 pp, low confidence
6Holo3-35B-A3B63.8%estimated ± 7.6 pp, medium confidence
7Claude Opus 4.859.9%measured
8Claude Opus 559.3%estimated ± 1.7 pp, low confidence
9Claude Fable 558.9%estimated ± 1.7 pp, low confidence
10Holo3-122B-A10B58.5%estimated ± 7.6 pp, medium confidence
11GPT-5.6 Sol58.0%measured
12Atria Dawn Preview57.3%estimated ± 3.5 pp, medium confidence
13Gemini 3.8 Flash56.7%estimated ± 1.7 pp, medium confidence
14GPT-5.5 Pro56.6%estimated ± 3.5 pp, high confidence
15Gemini 3.5 Flash56.5%measured
16GPT-5.4 Pro56.4%estimated ± 3.5 pp, high confidence
17Step 5 Preview56.2%estimated ± 3.5 pp, high confidence
18UI-Mate-27B56.0%estimated ± 7.6 pp, medium confidence
19Claude Mythos 555.9%estimated ± 3.5 pp, high confidence
20Claude Opus 5.555.9%estimated ± 6.3 pp, low confidence
21Claude Sonnet 5.555.9%estimated ± 6.3 pp, low confidence
22DeepSeek V4.1 Flash55.9%estimated ± 6.3 pp, medium confidence
23Gemini 4 Argon55.9%estimated ± 6.3 pp, low confidence
24GLM-5.355.9%estimated ± 6.3 pp, low confidence
25Grok 4.755.9%estimated ± 6.3 pp, low confidence
26Ling 3.1 Flash55.9%estimated ± 6.3 pp, low confidence
27MiMo-V2.6-Flash55.9%estimated ± 6.3 pp, medium confidence
28MiMo-V2.6-Pro55.9%estimated ± 6.3 pp, low confidence
29Qwen3.8-Flash-Next55.9%estimated ± 6.3 pp, low confidence
30Qwen3.8 Max Preview55.9%estimated ± 6.3 pp, low confidence
31GPT-6.1 Sol55.9%estimated ± 6.3 pp, medium confidence
32GPT-6 Luna55.9%estimated ± 6.3 pp, medium confidence
33GPT-6 Sol55.9%estimated ± 6.3 pp, medium confidence
34Mistral Large 455.9%estimated ± 6.3 pp, medium confidence
35Muse Spark 1.255.9%estimated ± 6.3 pp, medium confidence
36Qwen3.8-27B55.9%estimated ± 6.3 pp, medium confidence
37Grok 4.555.9%estimated ± 6.3 pp, medium confidence
38Gemini 3.6 Flash55.8%estimated ± 6.3 pp, medium confidence
39Hy3 Preview55.8%estimated ± 6.3 pp, medium confidence
40Apodex 1.155.8%estimated ± 6.3 pp, medium confidence
41Apodex 1.1 Mini55.8%estimated ± 6.3 pp, medium confidence
42Quasar 438B55.7%estimated ± 6.3 pp, medium confidence
43Ling 3.0 Flash VL55.7%estimated ± 6.3 pp, medium confidence
44GPT-5.555.6%measured
45Muse Spark 1.355.6%estimated ± 1.7 pp, medium confidence
46Ornith-1.5-397B55.4%estimated ± 3.5 pp, high confidence
47Kimi K355.4%estimated ± 1.7 pp, medium confidence
48Qwen3.7 Max55.4%estimated ± 6.3 pp, medium confidence
49Claude Mythos Preview54.9%estimated ± 9.0 pp, low confidence
50Claude Sonnet 554.9%estimated ± 1.7 pp, medium confidence
51Gemini 3.7 Flash54.9%estimated ± 1.7 pp, medium confidence
52GPT-5.454.6%measured
53Grok 4.654.3%estimated ± 1.7 pp, medium confidence
54MiniMax M354.0%estimated ± 3.5 pp, high confidence
55Grok 4.354.0%estimated ± 6.3 pp, medium confidence
56dots3-note Preview53.9%estimated ± 3.5 pp, high confidence
57GPT-5.6 Luna53.4%measured
58GPT-5.6 Terra53.1%measured
59Claude Opus 4.753.0%estimated ± 1.7 pp, medium confidence
60Hy352.8%estimated ± 6.3 pp, medium confidence
61Hy4 preview52.5%estimated ± 5.2 pp, low confidence
62Claude Opus 4.652.5%estimated ± 1.7 pp, low confidence
63MiMo-V2.552.3%estimated ± 6.9 pp, medium confidence
64Claude Sonnet 4.651.9%estimated ± 1.7 pp, low confidence
65DeepSeek V4 Pro 081351.8%measured
66Claude Opus 4.7 (Adaptive)51.7%estimated ± 3.5 pp, high confidence
67Inkling-Small50.4%estimated ± 3.5 pp, high confidence
68Beam50.4%estimated ± 3.5 pp, high confidence
69Claude Opus 4.6 (Adaptive)50.3%estimated ± 7.0 pp, low confidence
70Inkling50.2%estimated ± 3.5 pp, high confidence
71Gemini 3 Pro50.1%estimated ± 7.2 pp, medium confidence
72Kimi K2.7 Code50.1%estimated ± 6.3 pp, medium confidence
73Kimi K2.650.0%measured
74MiMo-V2-Pro49.7%estimated ± 6.9 pp, medium confidence
75Step 3.7 Flash49.5%measured
76Qwen3.8 Max49.4%estimated ± 5.2 pp, low confidence
77Agents-A148.9%estimated ± 3.5 pp, high confidence
78GLM-5.248.2%measured
79DeepSeek V4 Flash 073147.8%measured
80GPT-5.3 Codex46.6%estimated ± 7.2 pp, medium confidence
81MiniMax M2.746.3%measured
82Gemini 3 Flash46.1%estimated ± 7.2 pp, medium confidence
83Ling 3.0 Flash46.0%estimated ± 3.5 pp, high confidence
84MiMo-V2.5-Pro45.6%estimated ± 6.0 pp, low confidence
85UI-Mate-9B43.8%estimated ± 7.6 pp, medium confidence
86Claude Opus 4.543.5%measured
87Muse Spark43.3%estimated ± 6.3 pp, medium confidence
88GPT-5.2-Codex43.0%estimated ± 7.2 pp, medium confidence
89GPT-5.4 mini42.9%measured
90GPT-5.1-Codex41.6%estimated ± 7.2 pp, medium confidence
91GLM-5.141.4%estimated ± 3.5 pp, high confidence
92Grok Build 0.141.2%estimated ± 7.2 pp, medium confidence
93Ornith-1.5-35B-A3B40.9%estimated ± 3.5 pp, high confidence
94Claude Sonnet 4.540.8%estimated ± 7.2 pp, medium confidence
95Qwen3.6-27B40.4%estimated ± 6.3 pp, medium confidence
96Solar Open 240.3%estimated ± 7.5 pp, medium confidence
97Grok 4.1 Fast40.0%estimated ± 7.2 pp, medium confidence
98Gemini 3.5 Flash-Lite40.0%estimated ± 6.3 pp, medium confidence
99Agents-A1-4B39.9%estimated ± 3.5 pp, high confidence
100Qwen3.6 Plus39.8%measured
101GPT-5.238.6%estimated ± 3.5 pp, high confidence
102GLM-538.0%measured
103Qwen3 Max37.5%estimated ± 7.2 pp, medium confidence
104LLaDA2.2-flash37.2%estimated ± 7.5 pp, medium confidence
105Grok 436.6%estimated ± 7.2 pp, medium confidence
106Qwen3.5 397B36.3%measured
107Qwen3.5-122B-A10B35.9%estimated ± 3.5 pp, high confidence
108A.X K235.6%estimated ± 6.3 pp, medium confidence
109GPT-5.4 nano35.5%measured
110Claude 4 Sonnet34.7%estimated ± 7.2 pp, low confidence
111Gemini 3.1 Flash-Lite33.8%estimated ± 7.2 pp, low confidence
112Grok 4.2033.7%estimated ± 7.2 pp, low confidence
113Claude 4.1 Opus33.1%estimated ± 7.4 pp, low confidence
114Ling 3.0 Flash FP833.1%estimated ± 6.3 pp, medium confidence
115Pokee-Isaac 28B31.8%estimated ± 6.0 pp, low confidence
116Qwen3.5-27B31.8%estimated ± 3.5 pp, high confidence
117Qwen3.5-35B-A3B31.8%estimated ± 3.5 pp, high confidence
118Kimi K2.5 (Reasoning)31.2%estimated ± 3.5 pp, high confidence
119Qwen3.5 Plus30.9%estimated ± 7.4 pp, low confidence
120GPT-5 (high)30.0%estimated ± 6.3 pp, medium confidence
121Claude Haiku 4.529.7%estimated ± 7.4 pp, low confidence
122GLM-5V-Turbo28.0%estimated ± 7.2 pp, low confidence
123Kimi K2.527.8%measured
124GPT-5.127.2%estimated ± 6.3 pp, low confidence
125DeepSeek V3.227.0%estimated ± 7.2 pp, low confidence
126Gemini 3.1 Pro27.0%estimated ± 6.3 pp, low confidence
127Muse Glimmer 30B27.0%estimated ± 6.3 pp, low confidence
128Qwen3.7 Plus27.0%estimated ± 6.3 pp, low confidence
129MiniCPM5-2B26.9%estimated ± 6.3 pp, low confidence
130MiMo-V2-Flash26.9%estimated ± 6.3 pp, low confidence
131Gemma 4 31B26.9%estimated ± 6.3 pp, low confidence
132GPT-OSS 120B26.9%estimated ± 6.3 pp, low confidence
133Gemma 4 26B A4B26.9%estimated ± 6.3 pp, low confidence
134Ling 3.0 Tiny26.9%estimated ± 6.3 pp, low confidence
135Celeris-126.9%estimated ± 6.3 pp, low confidence
136Command A+26.9%estimated ± 6.3 pp, low confidence
137DeepSeek V326.9%estimated ± 6.3 pp, low confidence
138DeepSeek V3 032426.9%estimated ± 6.3 pp, low confidence
139Gemini 2.5 Pro26.9%estimated ± 6.3 pp, low confidence
140Gemma 3 27B26.9%estimated ± 6.3 pp, low confidence
141Gemma 4 12B26.9%estimated ± 6.3 pp, low confidence
142Gemma 4 E2B26.9%estimated ± 6.3 pp, low confidence
143Gemma 4 E4B26.9%estimated ± 6.3 pp, low confidence
144GPT-4.1 mini26.9%estimated ± 6.3 pp, low confidence
145GPT-4.1 nano26.9%estimated ± 6.3 pp, low confidence
146GPT-4o26.9%estimated ± 6.3 pp, low confidence
147GPT-4o mini26.9%estimated ± 6.3 pp, low confidence
148GPT-OSS 20B26.9%estimated ± 6.3 pp, low confidence
149K-Exaone26.9%estimated ± 6.3 pp, low confidence
150Ling 2.6 Flash26.9%estimated ± 6.3 pp, low confidence
151Llama 4 Maverick26.9%estimated ± 6.3 pp, low confidence
152Llama 4 Scout26.9%estimated ± 6.3 pp, low confidence
153Mistral Large 326.9%estimated ± 6.3 pp, low confidence
154Mistral Small 426.9%estimated ± 6.3 pp, low confidence
155Mistral Small 4 (Reasoning)26.9%estimated ± 6.3 pp, low confidence
156Nemotron 3 Nano 30B26.9%estimated ± 6.3 pp, low confidence
157Nemotron 3 Nano Omni 30B A3B26.9%estimated ± 6.3 pp, low confidence
158Nemotron 3 Super 100B26.9%estimated ± 6.3 pp, low confidence
159North Mini Code26.9%estimated ± 6.3 pp, low confidence
160Solar Pro 326.9%estimated ± 6.3 pp, low confidence
161Trinity-Large-Preview26.9%estimated ± 6.3 pp, low confidence
162Trinity-Large-Thinking26.9%estimated ± 6.3 pp, low confidence
163Ultravox v0.6 Llama 3.3 70B26.9%estimated ± 6.3 pp, low confidence
164Qwen3.6-35B-A3B26.9%measured
165Ornith-1.5-9B24.6%estimated ± 3.5 pp, medium confidence
166Granite 4.2 30B24.5%estimated ± 6.0 pp, low confidence
167GPT-4.123.9%estimated ± 7.2 pp, low confidence
168Granite 4.2 8B18.8%estimated ± 6.0 pp, low confidence
169GLM-4.717.9%estimated ± 3.5 pp, medium confidence
170Grok 4.117.3%estimated ± 6.9 pp, low confidence
171Solar Pro 413.9%estimated ± 3.5 pp, medium confidence
172LongCat-Flash-Lite-Sparse13.2%estimated ± 3.5 pp, medium confidence
173Nemotron 3 Ultra8.4%estimated ± 3.5 pp, medium confidence
174Granite 4.2 3B8.3%estimated ± 6.0 pp, low confidence
175LFM2.5-2.6B4.5%estimated ± 6.0 pp, low confidence
176Nemotron 3.5 Lightning 30B A3B NVFP43.0%estimated ± 3.5 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General