benchgap
Agentic · tools

AA AutomationBench leaderboard

As of 2026-10-07, the highest measured score on AA AutomationBench is 77.5% by Gemini 4 Argon. 135 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 4 Argon77.5%measured
2Claude Sonnet 5.571.8%measured
3Claude Opus 5.569.5%measured
4DeepSeek V4.1 Flash68.9%measured
5GPT-6 Astra68.5%measured
6Grok 4.666.7%measured
7GPT-5.6 Sol65.6%estimated ± 5.5 pp, low confidence
8Grok 4.765.6%measured
9GPT-6.1 Sol64.9%measured
10Qwen3.8 Max Preview64.7%estimated ± 7.0 pp, medium confidence
11Ling 3.1 Flash64.4%estimated ± 7.0 pp, medium confidence
12Claude Fable 564.3%estimated ± 7.0 pp, medium confidence
13GLM-5.264.3%estimated ± 5.5 pp, low confidence
14GPT-5.6 Terra62.7%estimated ± 5.5 pp, low confidence
15GLM-5.362.2%measured
16GPT-6 Sol61.6%measured
17Claude Opus 561.3%estimated ± 5.7 pp, low confidence
18Claude Mythos Preview61.1%estimated ± 5.5 pp, low confidence
19GLM-5.3-Flash60.4%measured
20Muse Spark 1.260.0%estimated ± 7.0 pp, medium confidence
21Gemini 3.8 Flash59.9%measured
22Mistral Large 459.9%measured
23GPT-5.6 Luna59.6%estimated ± 5.5 pp, low confidence
24Claude Fable 5.159.4%measured
25MiMo-V2.6-Pro58.6%measured
26Kimi K358.3%measured
27GPT-5.5 Pro58.0%estimated ± 5.7 pp, low confidence
28Muse Spark 1.357.9%measured
29Hy4 preview56.8%estimated ± 5.0 pp, medium confidence
30MiMo-V2.6-Flash56.7%estimated ± 5.0 pp, medium confidence
31Qwen3.8-Flash-Next55.2%estimated ± 5.0 pp, medium confidence
32Muse Spark 1.154.9%estimated ± 5.0 pp, medium confidence
33Qwen3.8 Max54.5%estimated ± 5.0 pp, medium confidence
34GPT-5.4 Pro53.9%estimated ± 5.7 pp, low confidence
35Atria Dawn Preview53.6%estimated ± 5.0 pp, medium confidence
36GPT-6 Luna53.2%measured
37Claude Opus 4.7 (Adaptive)52.4%estimated ± 5.0 pp, medium confidence
38GPT-5.551.5%estimated ± 5.0 pp, medium confidence
39Step 5 Preview51.0%measured
40GPT-5.450.4%estimated ± 5.0 pp, medium confidence
41Claude Sonnet 4.649.9%estimated ± 5.0 pp, medium confidence
42Claude Opus 4.649.8%estimated ± 5.0 pp, medium confidence
43Gemini 3.7 Flash49.7%estimated ± 7.0 pp, medium confidence
44Grok 4.549.3%estimated ± 7.0 pp, medium confidence
45GPT-5.249.2%estimated ± 5.0 pp, medium confidence
46GPT-5.3 Codex49.0%estimated ± 5.0 pp, medium confidence
47Claude Opus 4.548.6%estimated ± 5.0 pp, low confidence
48Qwen3.8-27B48.2%measured
49Claude Sonnet 4.547.3%estimated ± 5.0 pp, low confidence
50GPT-5.1-Codex46.9%estimated ± 5.0 pp, low confidence
51GPT-5.2-Codex46.8%estimated ± 5.0 pp, low confidence
52Claude Mythos 546.4%estimated ± 5.7 pp, low confidence
53Claude 4.1 Opus45.7%estimated ± 5.0 pp, low confidence
54Qwen3.5 Plus44.7%estimated ± 5.0 pp, low confidence
55Claude 4 Sonnet44.7%estimated ± 5.0 pp, low confidence
56Claude Haiku 4.544.0%estimated ± 5.0 pp, low confidence
57Gemini 3 Flash42.7%estimated ± 5.0 pp, low confidence
58Gemini 3 Pro42.7%estimated ± 5.0 pp, low confidence
59Kimi K2.542.0%estimated ± 5.0 pp, low confidence
60GPT-5 (high)41.9%estimated ± 5.0 pp, low confidence
61Claude Opus 4.741.7%estimated ± 7.0 pp, medium confidence
62Gemini 3.5 Flash39.8%estimated ± 7.0 pp, medium confidence
63Ornith-1.5-397B37.9%estimated ± 5.7 pp, low confidence
64Claude Haiku 5.535.4%measured
65Claude Sonnet 526.9%estimated ± 5.7 pp, low confidence
66Gemini 3.6 Flash26.2%estimated ± 7.0 pp, medium confidence
67Claude Opus 4.824.8%estimated ± 5.7 pp, low confidence
68MiniMax M321.3%measured
69DeepSeek V4 Pro 081320.6%estimated ± 5.7 pp, low confidence
70dots3-note Preview20.2%estimated ± 5.7 pp, low confidence
71Kimi K2.619.7%estimated ± 5.7 pp, low confidence
72Hy3 Preview13.7%estimated ± 7.0 pp, medium confidence
73Apodex 1.111.4%estimated ± 7.0 pp, medium confidence
74Apodex 1.1 Mini11.4%estimated ± 7.0 pp, medium confidence
75Quasar 438B10.2%estimated ± 7.0 pp, medium confidence
76Ling 3.0 Flash VL8.6%estimated ± 7.0 pp, medium confidence
77Qwen3.7 Max6.8%estimated ± 7.0 pp, medium confidence
78Muse Glimmer 30B6.8%measured
79MiMo-V2.5-Pro6.5%estimated ± 7.0 pp, medium confidence
80Inkling-Small5.9%estimated ± 5.7 pp, low confidence
81Beam5.9%estimated ± 5.7 pp, low confidence
82Grok 4.35.2%estimated ± 7.0 pp, medium confidence
83Inkling5.0%measured
84Hy34.9%estimated ± 7.0 pp, medium confidence
85Step 3.7 Flash4.6%estimated ± 5.7 pp, low confidence
86Kimi K2.7 Code4.5%estimated ± 7.0 pp, medium confidence
87Agents-A14.4%estimated ± 5.7 pp, low confidence
88GPT-5.4 mini4.3%estimated ± 7.0 pp, medium confidence
89MiniMax M2.74.3%estimated ± 7.0 pp, medium confidence
90Muse Spark4.2%estimated ± 7.0 pp, medium confidence
91Qwen3.6 Plus4.1%estimated ± 7.0 pp, medium confidence
92Qwen3.6-27B4.1%estimated ± 7.0 pp, medium confidence
93Gemini 3.5 Flash-Lite4.1%estimated ± 7.0 pp, medium confidence
94A.X K24.0%estimated ± 7.0 pp, medium confidence
95GPT-5.4 nano3.9%estimated ± 7.0 pp, medium confidence
96Ling 3.0 Flash FP83.9%estimated ± 7.0 pp, medium confidence
97Qwen3.6-35B-A3B3.8%estimated ± 7.0 pp, medium confidence
98GPT-5.13.8%estimated ± 7.0 pp, medium confidence
99Gemini 3.1 Pro3.8%estimated ± 7.0 pp, medium confidence
100Qwen3.7 Plus3.8%estimated ± 7.0 pp, low confidence
101Mistral Medium 3.5 128B3.8%estimated ± 7.0 pp, low confidence
102MiniCPM5-2B3.7%estimated ± 7.0 pp, low confidence
103MiMo-V2-Flash3.7%estimated ± 7.0 pp, low confidence
104Gemma 4 31B3.7%estimated ± 7.0 pp, low confidence
105GPT-OSS 120B3.7%estimated ± 7.0 pp, low confidence
106Granite 4.2 30B3.7%estimated ± 7.0 pp, low confidence
107Gemma 4 26B A4B3.7%estimated ± 7.0 pp, low confidence
108Ling 3.0 Tiny3.7%estimated ± 7.0 pp, low confidence
109Celeris-13.7%estimated ± 7.0 pp, low confidence
110Command A+3.7%estimated ± 7.0 pp, low confidence
111DeepSeek V33.7%estimated ± 7.0 pp, low confidence
112DeepSeek V3 03243.7%estimated ± 7.0 pp, low confidence
113Gemini 2.5 Pro3.7%estimated ± 7.0 pp, low confidence
114Gemma 3 27B3.7%estimated ± 7.0 pp, low confidence
115Gemma 4 12B3.7%estimated ± 7.0 pp, low confidence
116Gemma 4 E2B3.7%estimated ± 7.0 pp, low confidence
117Gemma 4 E4B3.7%estimated ± 7.0 pp, low confidence
118GPT-4.1 mini3.7%estimated ± 7.0 pp, low confidence
119GPT-4.1 nano3.7%estimated ± 7.0 pp, low confidence
120GPT-4o3.7%estimated ± 7.0 pp, low confidence
121GPT-4o mini3.7%estimated ± 7.0 pp, low confidence
122GPT-OSS 20B3.7%estimated ± 7.0 pp, low confidence
123Granite 4.2 3B3.7%estimated ± 7.0 pp, low confidence
124Granite 4.2 8B3.7%estimated ± 7.0 pp, low confidence
125K-Exaone3.7%estimated ± 7.0 pp, low confidence
126LFM2.5-2.6B3.7%estimated ± 7.0 pp, low confidence
127Ling 2.6 Flash3.7%estimated ± 7.0 pp, low confidence
128Llama 4 Maverick3.7%estimated ± 7.0 pp, low confidence
129Llama 4 Scout3.7%estimated ± 7.0 pp, low confidence
130Mercury 2.53.7%estimated ± 7.0 pp, low confidence
131Mistral Large 33.7%estimated ± 7.0 pp, low confidence
132Mistral Small 43.7%estimated ± 7.0 pp, low confidence
133Mistral Small 4 (Reasoning)3.7%estimated ± 7.0 pp, low confidence
134Nemotron 3 Nano 30B3.7%estimated ± 7.0 pp, low confidence
135Nemotron 3 Nano Omni 30B A3B3.7%estimated ± 7.0 pp, low confidence
136Nemotron 3 Super 100B3.7%estimated ± 7.0 pp, low confidence
137North Mini Code3.7%estimated ± 7.0 pp, low confidence
138Solar Pro 33.7%estimated ± 7.0 pp, low confidence
139Trinity-Large-Preview3.7%estimated ± 7.0 pp, low confidence
140Trinity-Large-Thinking3.7%estimated ± 7.0 pp, low confidence
141Ultravox v0.6 Llama 3.3 70B3.7%estimated ± 7.0 pp, low confidence
142DeepSeek V4 Flash 07313.5%estimated ± 5.7 pp, low confidence
143Ling 3.0 Flash3.2%estimated ± 5.7 pp, low confidence
144Nemotron 3 Ultra3.0%measured
145GLM-5.12.7%estimated ± 5.7 pp, low confidence
146Ornith-1.5-35B-A3B2.7%estimated ± 5.7 pp, low confidence
147Agents-A1-4B2.7%estimated ± 5.7 pp, low confidence
148Qwen3.5-122B-A10B2.6%estimated ± 5.7 pp, low confidence
149Qwen3.5 397B2.6%estimated ± 5.7 pp, low confidence
150Qwen3.5-27B2.6%estimated ± 5.7 pp, low confidence
151Qwen3.5-35B-A3B2.6%estimated ± 5.7 pp, low confidence
152Kimi K2.5 (Reasoning)2.6%estimated ± 5.7 pp, low confidence
153Ornith-1.5-9B2.6%estimated ± 5.7 pp, low confidence
154GLM-4.72.6%estimated ± 5.7 pp, low confidence
155Solar Pro 42.6%estimated ± 5.7 pp, low confidence
156LongCat-Flash-Lite-Sparse2.6%estimated ± 5.7 pp, low confidence
157Nemotron 3.5 Lightning 30B A3B NVFP42.6%estimated ± 5.7 pp, low confidence
158GLM-50.0%estimated ± 13.0 pp, low confidence
159LLaDA2.2-flash0.0%estimated ± 13.0 pp, low confidence
160Solar Open 20.0%estimated ± 13.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General