benchgap
Agentic · tools

AutomationBench leaderboard

As of 2026-10-07, the highest measured score on AutomationBench is 54.8% by DeepSeek V4.1 Flash. 128 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1DeepSeek V4.1 Flash54.8%measured
2Atria Dawn Preview53.8%measured
3Nemotron 3 Ultra53.7%estimated ± 7.0 pp, low confidence
4Fugu Cyber53.2%estimated ± 5.8 pp, medium confidence
5MiMo-V2.6-Pro53.1%measured
6Gemini 3.8 Flash Cyber52.9%estimated ± 5.8 pp, medium confidence
7Ling 3.1 Flash52.5%measured
8MiMo-V2.6-Flash52.3%measured
9Gemini 4 Argon51.3%measured
10MiniMax M350.6%estimated ± 7.0 pp, low confidence
11Muse Spark 1.349.4%measured
12GLM-5.3-Flash48.8%measured
13GLM-5.348.2%measured
14Muse Glimmer 30B48.1%estimated ± 7.0 pp, low confidence
15Mistral Large 445.5%estimated ± 7.6 pp, low confidence
16GPT-5.6 Sol45.2%estimated ± 5.8 pp, medium confidence
17Inkling44.9%estimated ± 7.0 pp, low confidence
18Qwen3.8 Max Preview44.6%estimated ± 9.7 pp, low confidence
19Grok 4.744.6%estimated ± 7.6 pp, low confidence
20Qwen3.8-Flash-Next44.5%estimated ± 9.7 pp, low confidence
21Step 5 Preview44.0%measured
22Claude Haiku 5.544.0%estimated ± 7.6 pp, low confidence
23Gemini 3.8 Flash43.8%estimated ± 7.6 pp, low confidence
24GPT-6 Luna42.6%estimated ± 7.6 pp, low confidence
25GPT-6 Astra41.4%measured
26Claude Sonnet 5.540.4%estimated ± 7.6 pp, low confidence
27Claude Opus 5.540.0%measured
28Qwen3.8-27B38.9%estimated ± 7.0 pp, low confidence
29GPT-5.5 Pro38.4%estimated ± 10.3 pp, low confidence
30GPT-5.4 Pro37.5%estimated ± 10.3 pp, low confidence
31Claude Mythos 536.9%estimated ± 5.8 pp, medium confidence
32GPT-6.1 Sol36.1%measured
33Grok 4.635.0%estimated ± 7.0 pp, low confidence
34Ornith-1.5-397B34.8%estimated ± 10.3 pp, low confidence
35Gemini 3.5 Flash33.2%estimated ± 7.0 pp, low confidence
36GPT-6 Sol33.2%measured
37Claude Fable 532.3%estimated ± 7.0 pp, low confidence
38Hy4 preview32.1%measured
39Gemini 3.5 Flash Cyber31.8%estimated ± 5.8 pp, medium confidence
40DeepSeek V4 Pro 081331.8%measured
41dots3-note Preview31.8%estimated ± 10.3 pp, low confidence
42Claude Fable 5.131.4%measured
43Claude Mythos Preview31.3%estimated ± 5.8 pp, medium confidence
44Kimi K330.8%measured
45Gemini 3.7 Flash30.4%measured
46Muse Spark 1.229.0%estimated ± 9.7 pp, low confidence
47Claude Sonnet 528.7%estimated ± 9.7 pp, low confidence
48GPT-5.528.7%estimated ± 5.8 pp, medium confidence
49GPT-5.6 Terra28.7%estimated ± 5.8 pp, medium confidence
50Claude Opus 4.828.6%estimated ± 9.7 pp, low confidence
51Grok 4.528.4%estimated ± 9.7 pp, low confidence
52GLM-5.228.4%estimated ± 9.7 pp, low confidence
53Gemini 3.6 Flash28.4%estimated ± 9.7 pp, low confidence
54A.X K228.4%estimated ± 9.7 pp, low confidence
55Apodex 1.128.4%estimated ± 9.7 pp, low confidence
56Apodex 1.1 Mini28.4%estimated ± 9.7 pp, low confidence
57Celeris-128.4%estimated ± 9.7 pp, low confidence
58Command A+28.4%estimated ± 9.7 pp, low confidence
59DeepSeek V328.4%estimated ± 9.7 pp, low confidence
60DeepSeek V3 032428.4%estimated ± 9.7 pp, low confidence
61Gemini 2.5 Pro28.4%estimated ± 9.7 pp, low confidence
62Gemini 3.1 Pro28.4%estimated ± 9.7 pp, low confidence
63Gemini 3.5 Flash-Lite28.4%estimated ± 9.7 pp, low confidence
64Gemma 3 27B28.4%estimated ± 9.7 pp, low confidence
65Gemma 4 12B28.4%estimated ± 9.7 pp, low confidence
66Gemma 4 26B A4B28.4%estimated ± 9.7 pp, low confidence
67Gemma 4 31B28.4%estimated ± 9.7 pp, low confidence
68Gemma 4 E2B28.4%estimated ± 9.7 pp, low confidence
69Gemma 4 E4B28.4%estimated ± 9.7 pp, low confidence
70GLM-4.728.4%estimated ± 9.7 pp, low confidence
71GPT-4.1 mini28.4%estimated ± 9.7 pp, low confidence
72GPT-4.1 nano28.4%estimated ± 9.7 pp, low confidence
73GPT-4o28.4%estimated ± 9.7 pp, low confidence
74GPT-4o mini28.4%estimated ± 9.7 pp, low confidence
75GPT-5.128.4%estimated ± 9.7 pp, low confidence
76GPT-5.4 mini28.4%estimated ± 9.7 pp, low confidence
77GPT-5.4 nano28.4%estimated ± 9.7 pp, low confidence
78GPT-5 (high)28.4%estimated ± 9.7 pp, low confidence
79GPT-OSS 120B28.4%estimated ± 9.7 pp, low confidence
80GPT-OSS 20B28.4%estimated ± 9.7 pp, low confidence
81Granite 4.2 30B28.4%estimated ± 9.7 pp, low confidence
82Granite 4.2 3B28.4%estimated ± 9.7 pp, low confidence
83Granite 4.2 8B28.4%estimated ± 9.7 pp, low confidence
84Grok 4.328.4%estimated ± 9.7 pp, low confidence
85Hy328.4%estimated ± 9.7 pp, low confidence
86Hy3 Preview28.4%estimated ± 9.7 pp, low confidence
87Inkling-Small28.4%estimated ± 9.7 pp, low confidence
88K-Exaone28.4%estimated ± 9.7 pp, low confidence
89Kimi K2.628.4%estimated ± 9.7 pp, low confidence
90Kimi K2.528.4%estimated ± 9.7 pp, low confidence
91Kimi K2.5 (Reasoning)28.4%estimated ± 9.7 pp, low confidence
92Kimi K2.7 Code28.4%estimated ± 9.7 pp, low confidence
93LFM2.5-2.6B28.4%estimated ± 9.7 pp, low confidence
94Ling 2.6 Flash28.4%estimated ± 9.7 pp, low confidence
95Ling 3.0 Flash28.4%estimated ± 9.7 pp, low confidence
96Ling 3.0 Flash FP828.4%estimated ± 9.7 pp, low confidence
97Ling 3.0 Flash VL28.4%estimated ± 9.7 pp, low confidence
98Ling 3.0 Tiny28.4%estimated ± 9.7 pp, low confidence
99Llama 4 Maverick28.4%estimated ± 9.7 pp, low confidence
100Llama 4 Scout28.4%estimated ± 9.7 pp, low confidence
101Mercury 2.528.4%estimated ± 9.7 pp, low confidence
102MiMo-V2.5-Pro28.4%estimated ± 9.7 pp, low confidence
103MiMo-V2-Flash28.4%estimated ± 9.7 pp, low confidence
104MiniCPM5-2B28.4%estimated ± 9.7 pp, low confidence
105MiniMax M2.728.4%estimated ± 9.7 pp, low confidence
106Mistral Large 328.4%estimated ± 9.7 pp, low confidence
107Mistral Medium 3.5 128B28.4%estimated ± 9.7 pp, low confidence
108Mistral Small 428.4%estimated ± 9.7 pp, low confidence
109Mistral Small 4 (Reasoning)28.4%estimated ± 9.7 pp, low confidence
110Nemotron 3.5 Lightning 30B A3B NVFP428.4%estimated ± 9.7 pp, low confidence
111Nemotron 3 Nano 30B28.4%estimated ± 9.7 pp, low confidence
112Nemotron 3 Nano Omni 30B A3B28.4%estimated ± 9.7 pp, low confidence
113Nemotron 3 Super 100B28.4%estimated ± 9.7 pp, low confidence
114North Mini Code28.4%estimated ± 9.7 pp, low confidence
115Quasar 438B28.4%estimated ± 9.7 pp, low confidence
116Qwen3.5-122B-A10B28.4%estimated ± 9.7 pp, low confidence
117Qwen3.6-27B28.4%estimated ± 9.7 pp, low confidence
118Qwen3.6-35B-A3B28.4%estimated ± 9.7 pp, low confidence
119Qwen3.6 Plus28.4%estimated ± 9.7 pp, low confidence
120Qwen3.7 Max28.4%estimated ± 9.7 pp, low confidence
121Qwen3.7 Plus28.4%estimated ± 9.7 pp, low confidence
122Solar Pro 328.4%estimated ± 9.7 pp, low confidence
123Solar Pro 428.4%estimated ± 9.7 pp, low confidence
124Step 3.7 Flash28.4%estimated ± 9.7 pp, low confidence
125Trinity-Large-Preview28.4%estimated ± 9.7 pp, low confidence
126Trinity-Large-Thinking28.4%estimated ± 9.7 pp, low confidence
127Ultravox v0.6 Llama 3.3 70B28.4%estimated ± 9.7 pp, low confidence
128GPT-5.428.4%estimated ± 5.8 pp, medium confidence
129GPT-5.6 Luna28.4%estimated ± 5.8 pp, medium confidence
130Claude Opus 4.528.4%estimated ± 5.8 pp, low confidence
131Claude Opus 4.628.4%estimated ± 5.8 pp, low confidence
132Claude Opus 4.7 (Adaptive)28.4%estimated ± 5.8 pp, low confidence
133Claude Sonnet 4.628.4%estimated ± 5.8 pp, low confidence
134GLM-528.4%estimated ± 5.8 pp, low confidence
135GLM-5.128.4%estimated ± 5.8 pp, low confidence
136Muse Spark28.4%estimated ± 5.8 pp, low confidence
137Muse Spark 1.128.4%estimated ± 5.8 pp, low confidence
138Qwen3.8 Max27.3%measured
139Beam27.1%estimated ± 10.3 pp, low confidence
140Claude Opus 526.0%measured
141Agents-A125.7%estimated ± 10.3 pp, low confidence
142DeepSeek V4 Flash 073125.1%measured
143Ornith-1.5-35B-A3B20.8%estimated ± 10.3 pp, low confidence
144Agents-A1-4B20.3%estimated ± 10.3 pp, low confidence
145GPT-5.219.8%estimated ± 10.3 pp, low confidence
146Qwen3.5 397B17.8%estimated ± 10.3 pp, low confidence
147Qwen3.5-27B17.3%estimated ± 10.3 pp, low confidence
148Qwen3.5-35B-A3B17.3%estimated ± 10.3 pp, low confidence
149Ornith-1.5-9B15.2%estimated ± 10.3 pp, low confidence
150LongCat-Flash-Lite-Sparse12.1%estimated ± 10.3 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General