Agentic · tools
AutomationBench leaderboard
As of 2026-10-07, the highest measured score on AutomationBench is 77.5% by Gemini 4 Argon (high). 11 more models have estimated scores, calibrated from the benchmarks they were measured on.
Measured scores: artificial-analysis.
| # | Model | Score | Source |
|---|---|---|---|
| 1 | Gemini 3.7 Flash (high) | 78.6% | estimated ± 11.2 pp, low confidence |
| 2 | Gemini 4 Argon (high) | 77.5% | measured |
| 3 | Claude Sonnet 5.5 (max with fallback) | 71.8% | measured |
| 4 | Claude Opus 5.5 (max with fallback) | 69.5% | measured |
| 5 | DeepSeek V4.1 Flash (max) | 68.9% | measured |
| 6 | GPT-6 Astra (max) | 68.5% | measured |
| 7 | Claude Opus 5 (max) | 67.5% | estimated ± 11.2 pp, low confidence |
| 8 | GPT-6 Astra (xhigh) | 67.2% | measured |
| 9 | Grok 4.6 (xhigh) | 67.0% | measured |
| 10 | Grok 4.6 (high) | 66.7% | measured |
| 11 | GPT-6 Astra (high) | 66.6% | measured |
| 12 | GPT-6.1 Sol (xhigh) | 66.6% | measured |
| 13 | Grok 4.7 (xhigh) | 65.6% | measured |
| 14 | GPT-6.1 Sol (max) | 64.9% | measured |
| 15 | Qwen3.8 Max | 62.8% | estimated ± 10.6 pp, low confidence |
| 16 | GLM-5.3 (max) | 62.2% | measured |
| 17 | GPT-6 Sol (max) | 61.6% | measured |
| 18 | Qwen3.8 2.4T A95B | 61.2% | estimated ± 10.6 pp, low confidence |
| 19 | GPT-5.5 (xhigh) | 61.2% | estimated ± 11.2 pp, low confidence |
| 20 | GLM-5.3-Flash | 60.4% | measured |
| 21 | Gemini 3.8 Flash (high) | 59.9% | measured |
| 22 | Mistral Large 4 Preview | 59.9% | measured |
| 23 | Claude Fable 5.1 (xhigh with fallback) | 59.6% | estimated ± 13.0 pp, low confidence |
| 24 | Qwen3.8 27B (medium) | 59.6% | estimated ± 10.6 pp, low confidence |
| 25 | Claude Fable 5.1 (max with fallback) | 59.4% | measured |
| 26 | Muse Spark 1.3 (xhigh) | 59.4% | estimated ± 10.6 pp, low confidence |
| 27 | Claude Fable 5 (with fallback) | 59.2% | estimated ± 11.2 pp, low confidence |
| 28 | MiMo-V2.6-Pro | 58.6% | measured |
| 29 | Kimi K3 (max) | 58.3% | measured |
| 30 | Claude Fable 5.1 (high with fallback) | 58.1% | estimated ± 13.0 pp, low confidence |
| 31 | Muse Spark 1.3 (max) | 57.9% | measured |
| 32 | GPT-5.6 Sol (max) | 57.2% | estimated ± 11.2 pp, low confidence |
| 33 | Qwen3.8 Max (0902) | 56.2% | measured |
| 34 | GPT-6 Luna (max) | 53.2% | measured |
| 35 | Step 5 Preview | 51.0% | measured |
| 36 | Qwen3.8 27B (xhigh) | 48.2% | measured |
| 37 | K2 Horizon 375B A23B | 37.2% | measured |
| 38 | MiniMax-M3 | 21.3% | measured |
| 39 | Muse Glimmer (high) | 6.8% | measured |
| 40 | Inkling (xhigh) | 5.0% | measured |
| 41 | Nemotron 3 Ultra | 3.0% | measured |
All leaderboards
Agentic · terminal
Agentic · tools
- AA-AnalystAgent
- APEX-Agents
- AutomationBench
- EnterpriseOps-Gym
- Harvey LAB
- ITBench SRE
- τ³-Bench Banking
- τ²-Bench Telecom