benchgap
Agentic · tools

AutomationBench leaderboard

As of 2026-10-07, the highest measured score on AutomationBench is 77.5% by Gemini 4 Argon (high). 11 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Gemini 3.7 Flash (high)78.6%estimated ± 11.2 pp, low confidence
2Gemini 4 Argon (high)77.5%measured
3Claude Sonnet 5.5 (max with fallback)71.8%measured
4Claude Opus 5.5 (max with fallback)69.5%measured
5DeepSeek V4.1 Flash (max)68.9%measured
6GPT-6 Astra (max)68.5%measured
7Claude Opus 5 (max)67.5%estimated ± 11.2 pp, low confidence
8GPT-6 Astra (xhigh)67.2%measured
9Grok 4.6 (xhigh)67.0%measured
10Grok 4.6 (high)66.7%measured
11GPT-6 Astra (high)66.6%measured
12GPT-6.1 Sol (xhigh)66.6%measured
13Grok 4.7 (xhigh)65.6%measured
14GPT-6.1 Sol (max)64.9%measured
15Qwen3.8 Max62.8%estimated ± 10.6 pp, low confidence
16GLM-5.3 (max)62.2%measured
17GPT-6 Sol (max)61.6%measured
18Qwen3.8 2.4T A95B61.2%estimated ± 10.6 pp, low confidence
19GPT-5.5 (xhigh)61.2%estimated ± 11.2 pp, low confidence
20GLM-5.3-Flash60.4%measured
21Gemini 3.8 Flash (high)59.9%measured
22Mistral Large 4 Preview59.9%measured
23Claude Fable 5.1 (xhigh with fallback)59.6%estimated ± 13.0 pp, low confidence
24Qwen3.8 27B (medium)59.6%estimated ± 10.6 pp, low confidence
25Claude Fable 5.1 (max with fallback)59.4%measured
26Muse Spark 1.3 (xhigh)59.4%estimated ± 10.6 pp, low confidence
27Claude Fable 5 (with fallback)59.2%estimated ± 11.2 pp, low confidence
28MiMo-V2.6-Pro58.6%measured
29Kimi K3 (max)58.3%measured
30Claude Fable 5.1 (high with fallback)58.1%estimated ± 13.0 pp, low confidence
31Muse Spark 1.3 (max)57.9%measured
32GPT-5.6 Sol (max)57.2%estimated ± 11.2 pp, low confidence
33Qwen3.8 Max (0902)56.2%measured
34GPT-6 Luna (max)53.2%measured
35Step 5 Preview51.0%measured
36Qwen3.8 27B (xhigh)48.2%measured
37K2 Horizon 375B A23B37.2%measured
38MiniMax-M321.3%measured
39Muse Glimmer (high)6.8%measured
40Inkling (xhigh)5.0%measured
41Nemotron 3 Ultra3.0%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents