benchgap
Agentic · tools

AA Agentic Index leaderboard

As of 2026-10-07, the highest measured score on AA Agentic Index is 58.0% by Claude Fable 5.1. 151 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.561.3%estimated ± 2.8 pp, medium confidence
2Claude Sonnet 5.560.3%estimated ± 2.8 pp, medium confidence
3Claude Fable 5.158.0%measured
4GLM-5.3-Flash56.2%estimated ± 3.3 pp, medium confidence
5Claude Opus 556.2%measured
6MiMo-V2.6-Pro56.0%estimated ± 1.7 pp, high confidence
7Muse Spark 1.355.7%measured
8Grok 4.755.1%estimated ± 2.8 pp, high confidence
9GLM-5.353.4%measured
10Grok 4.653.4%measured
11Fugu Cyber53.1%estimated ± 5.8 pp, low confidence
12MiMo-V2.6-Flash52.7%estimated ± 1.7 pp, high confidence
13Atria Dawn Preview52.5%estimated ± 5.8 pp, low confidence
14Ling 3.1 Flash52.0%estimated ± 1.7 pp, high confidence
15Gemini 3.8 Flash Cyber52.0%estimated ± 5.8 pp, low confidence
16GPT-6 Astra51.5%measured
17Qwen3.8-Flash-Next51.4%estimated ± 2.8 pp, high confidence
18Gemini 4 Argon51.1%estimated ± 2.8 pp, high confidence
19DeepSeek V4.1 Flash51.0%estimated ± 1.7 pp, high confidence
20Claude Fable 551.0%measured
21Kimi K350.6%measured
22Hy4 preview50.6%estimated ± 3.3 pp, medium confidence
23GPT-5.6 Sol50.5%measured
24Step 5 Preview49.8%estimated ± 1.7 pp, high confidence
25Qwen3.8 Max Preview49.6%measured
26DeepSeek V4 Pro 081349.6%measured
27GPT-5.5 Pro49.1%estimated ± 8.2 pp, medium confidence
28GPT-6.1 Sol48.8%estimated ± 2.8 pp, high confidence
29Qwen3.8 Max48.5%estimated ± 3.3 pp, medium confidence
30Claude Mythos 548.5%estimated ± 5.8 pp, medium confidence
31Gemini 3.5 Flash Cyber47.7%estimated ± 5.8 pp, medium confidence
32GPT-5.4 Pro47.6%estimated ± 8.2 pp, medium confidence
33Claude Mythos Preview47.5%estimated ± 5.8 pp, medium confidence
34Ornith-1.5-397B46.8%estimated ± 3.3 pp, medium confidence
35Qwen3.8-27B46.5%measured
36GPT-6 Sol45.6%estimated ± 2.8 pp, high confidence
37Claude Sonnet 544.3%measured
38Muse Spark 1.244.0%measured
39GPT-5.6 Terra43.7%measured
40GPT-5.6 Luna42.7%measured
41Claude Opus 4.842.6%measured
42Grok 4.542.1%measured
43GPT-6 Luna42.1%estimated ± 2.8 pp, high confidence
44DeepSeek V4 Flash 073141.7%measured
45Mistral Large 441.4%estimated ± 2.8 pp, high confidence
46Gemini 3.8 Flash41.1%measured
47Holo3-35B-A3B39.8%estimated ± 8.0 pp, medium confidence
48Claude Opus 4.7 (Adaptive)39.5%measured
49GLM-5.239.4%measured
50GPT-5.537.3%measured
51Claude Haiku 5.537.1%estimated ± 6.7 pp, medium confidence
52Gemini 3.7 Flash36.4%measured
53GPT-5.433.2%estimated ± 1.7 pp, high confidence
54Claude Opus 4.6 (Adaptive)32.9%estimated ± 11.8 pp, low confidence
55Quasar 438B32.7%measured
56Holo3-122B-A10B31.5%estimated ± 8.0 pp, medium confidence
57Beam31.1%estimated ± 8.2 pp, medium confidence
58MiniMax M330.8%measured
59Gemini 3.6 Flash30.2%measured
60Apodex 1.129.4%estimated ± 2.8 pp, high confidence
61Apodex 1.1 Mini29.4%estimated ± 2.8 pp, high confidence
62Agents-A129.2%estimated ± 8.2 pp, medium confidence
63Claude Opus 4.728.7%estimated ± 6.4 pp, medium confidence
64UI-Mate-27B28.3%estimated ± 8.0 pp, medium confidence
65K-EXAONE 2.027.7%estimated ± 5.1 pp, low confidence
66Ling 3.0 Flash VL27.6%estimated ± 2.8 pp, high confidence
67Muse Spark 1.127.5%measured
68Ornith-1.0-397B27.4%estimated ± 5.1 pp, low confidence
69Gemini 3.5 Flash27.3%measured
70dots3-note Preview26.5%estimated ± 3.3 pp, medium confidence
71Gemini 3 Pro26.2%estimated ± 6.4 pp, medium confidence
72Hy325.6%measured
73Hy3 Preview25.6%measured
74GLM-5.125.2%measured
75Inkling-Small25.0%measured
76Qwen 3.6 Max (preview)24.8%estimated ± 12.3 pp, low confidence
77Solar Pro 424.7%estimated ± 2.8 pp, high confidence
78Inkling24.3%measured
79Qwen3.7 Max23.9%measured
80Grok 4.1 Fast (Reasoning)23.9%estimated ± 12.3 pp, low confidence
81Claude Opus 4.623.7%estimated ± 5.1 pp, medium confidence
82Ornith-1.0-35B23.4%estimated ± 5.1 pp, medium confidence
83MiMo-V2.5-Pro22.7%measured
84Kimi K2.7 Code22.5%measured
85Claude Opus 4.5 Thinking22.5%estimated ± 12.3 pp, low confidence
86Claude Sonnet 4.622.4%estimated ± 5.1 pp, medium confidence
87Kimi K2.622.1%measured
88Agents-A1-4B22.0%estimated ± 8.2 pp, medium confidence
89LFM2.5-8B-A1B22.0%estimated ± 12.3 pp, low confidence
90Nemotron 3 Ultra21.7%measured
91GPT-5 (medium)21.4%estimated ± 12.3 pp, low confidence
92GPT-5.3 Codex21.1%estimated ± 6.4 pp, medium confidence
93Ling 3.0 Flash21.0%measured
94Ling 3.0 Flash FP821.0%measured
95LLaDA2.2-flash20.6%estimated ± 5.1 pp, medium confidence
96Qwen3.5 397B (Reasoning)20.4%estimated ± 12.3 pp, low confidence
97Ornith-1.0-9B20.1%estimated ± 5.1 pp, medium confidence
98GPT-5.1-Codex-Max20.1%estimated ± 12.3 pp, low confidence
99Qwen3.6-27B20.1%measured
100GLM-4.719.8%estimated ± 2.8 pp, high confidence
101Step 3.7 Flash19.8%estimated ± 2.8 pp, high confidence
102MiMo-V2.519.7%estimated ± 5.1 pp, medium confidence
103Qwen3.7 Plus19.7%measured
104GPT-5.4 mini19.7%measured
105o319.2%estimated ± 12.3 pp, low confidence
106Ternary Bonsai 2 27B19.1%estimated ± 12.3 pp, low confidence
107Laguna S 2.118.8%estimated ± 3.3 pp, low confidence
108Qwen3.6 Plus18.7%estimated ± 2.8 pp, high confidence
109Claude Opus 4.518.5%estimated ± 5.1 pp, medium confidence
110GLM-4.617.8%estimated ± 12.3 pp, low confidence
111GPT-5.4 nano17.7%measured
112MiMo-V2-Pro17.7%estimated ± 5.1 pp, medium confidence
113GLM-517.7%estimated ± 5.1 pp, medium confidence
114Ornith-1.5-35B-A3B17.5%estimated ± 3.3 pp, low confidence
115LLaDA2.2-mini17.5%estimated ± 5.1 pp, medium confidence
116Grok Code Fast 117.4%estimated ± 12.3 pp, low confidence
117Qwen3.5 397B17.3%estimated ± 5.1 pp, medium confidence
118Grok 4.317.2%measured
119A.X K217.2%estimated ± 2.8 pp, high confidence
120GPT-5.2-Codex17.0%estimated ± 6.4 pp, medium confidence
121GLM-5-Turbo16.9%estimated ± 5.1 pp, medium confidence
122MiniMax M2.716.8%measured
123UI-Mate-9B16.7%estimated ± 8.0 pp, medium confidence
124GLM-5V-Turbo16.1%estimated ± 5.1 pp, medium confidence
125Gemini 3.5 Flash-Lite15.9%measured
126Claude 4.1 Opus Thinking15.8%estimated ± 12.3 pp, low confidence
127Muse Spark15.8%estimated ± 1.7 pp, high confidence
128GPT-5.1-Codex15.7%estimated ± 6.4 pp, medium confidence
129Grok 4.115.5%estimated ± 6.8 pp, medium confidence
130Grok Build 0.115.4%estimated ± 6.4 pp, medium confidence
131Claude Sonnet 4.515.1%estimated ± 6.4 pp, medium confidence
132GPT-5 (high)15.0%estimated ± 2.8 pp, high confidence
133Qwen3.6-35B-A3B15.0%measured
134Grok 4.1 Fast14.4%estimated ± 6.4 pp, medium confidence
135Gemini 3 Flash14.2%estimated ± 5.1 pp, medium confidence
136GPT-5.214.0%estimated ± 6.4 pp, medium confidence
137Grok 4 Fast (Reasoning)13.8%estimated ± 12.3 pp, low confidence
138MiMo-V2-Omni12.7%estimated ± 5.1 pp, medium confidence
139o112.6%estimated ± 12.3 pp, low confidence
140Qwen3 Max12.6%estimated ± 6.4 pp, medium confidence
141LongCat-Flash-Lite-Sparse12.2%estimated ± 8.2 pp, medium confidence
142Kimi K212.0%estimated ± 12.3 pp, low confidence
143Grok 411.9%estimated ± 6.4 pp, medium confidence
144Kimi K2.511.6%estimated ± 2.8 pp, high confidence
145Kimi K2.5 (Reasoning)11.6%estimated ± 2.8 pp, high confidence
146GPT-5.111.0%estimated ± 2.8 pp, high confidence
147DeepSeek V3.211.0%estimated ± 5.1 pp, medium confidence
148Claude 4 Sonnet10.7%estimated ± 6.4 pp, medium confidence
149Qwen3.5-27B10.6%estimated ± 6.4 pp, medium confidence
150Muse Glimmer 30B10.5%measured
151Gemini 3.1 Pro10.3%measured
152Gemini 3.1 Flash-Lite10.2%estimated ± 6.4 pp, medium confidence
153Grok 4.2010.2%estimated ± 6.4 pp, medium confidence
154Qwen3.5-122B-A10B9.6%measured
155Mistral Medium 3.5 128B9.4%measured
156Ornith-1.5-9B7.7%estimated ± 3.3 pp, low confidence
157Sarvam 105B6.8%estimated ± 12.3 pp, low confidence
158Qwen3.5-35B-A3B6.8%estimated ± 6.4 pp, low confidence
159Gemma 4 31B6.7%measured
160GLM-4.5-Air6.7%estimated ± 12.3 pp, low confidence
161MiniCPM5-2B6.5%estimated ± 2.8 pp, high confidence
162GPT-OSS 120B6.2%measured
163Nemotron 3.5 Lightning 30B A3B NVFP46.1%measured
164GPT-4.15.7%estimated ± 6.4 pp, low confidence
165MiMo-V2-Flash4.4%estimated ± 2.8 pp, high confidence
166Nemotron 3 Super 100B4.1%measured
167Granite 4.2 8B3.7%measured
168Command A+3.6%measured
169Gemini 2.5 Pro3.6%measured
170Granite 4.2 30B3.3%estimated ± 2.8 pp, high confidence
171DeepSeek V3.1 (Reasoning)3.3%estimated ± 12.3 pp, low confidence
172Gemma 4 26B A4B3.2%estimated ± 2.8 pp, high confidence
173Ling 3.0 Tiny3.1%estimated ± 2.8 pp, high confidence
174DeepSeek-R13.0%estimated ± 12.3 pp, low confidence
175DeepSeek V3 03242.7%estimated ± 2.8 pp, high confidence
176Gemma 4 12B2.7%estimated ± 2.8 pp, high confidence
177Gemma 4 E2B2.7%estimated ± 2.8 pp, high confidence
178Gemma 4 E4B2.7%estimated ± 2.8 pp, high confidence
179GPT-4.1 mini2.7%estimated ± 2.8 pp, high confidence
180GPT-4.1 nano2.7%estimated ± 2.8 pp, high confidence
181GPT-4o2.7%estimated ± 2.8 pp, high confidence
182GPT-4o mini2.7%estimated ± 2.8 pp, high confidence
183Granite 4.2 3B2.7%estimated ± 2.8 pp, high confidence
184K-Exaone2.7%estimated ± 2.8 pp, high confidence
185LFM2.5-2.6B2.7%estimated ± 2.8 pp, high confidence
186Ling 2.6 Flash2.7%estimated ± 2.8 pp, high confidence
187Mercury 2.52.7%estimated ± 2.8 pp, high confidence
188Nemotron 3 Nano Omni 30B A3B2.7%estimated ± 2.8 pp, high confidence
189North Mini Code2.7%estimated ± 2.8 pp, high confidence
190Solar Pro 32.7%estimated ± 2.8 pp, high confidence
191Ultravox v0.6 Llama 3.3 70B2.7%estimated ± 2.8 pp, high confidence
192Mistral Large 32.4%measured
193DeepSeek V3.12.4%estimated ± 12.3 pp, low confidence
194Sarvam 30B2.3%estimated ± 12.3 pp, low confidence
195Mistral Small 41.4%measured
196Mistral Small 4 (Reasoning)1.4%measured
197GPT-OSS 20B1.4%measured
198Solar Pro 21.3%estimated ± 12.3 pp, low confidence
199Trinity-Large-Preview1.2%measured
200Trinity-Large-Thinking1.2%measured
201Nemotron 3 Nano 30B1.0%measured
202Mistral Large 20.9%estimated ± 12.3 pp, low confidence
203DeepSeek V30.8%measured
204Celeris-10.7%measured
205Llama 4 Maverick0.6%measured
206Llama 4 Scout0.5%measured
207Gemma 3 27B0.1%measured
208o3-mini0.1%estimated ± 12.3 pp, low confidence
209Claude 3 Haiku0.0%estimated ± 12.3 pp, low confidence
210Exaone 4.0 1.2B0.0%estimated ± 12.3 pp, low confidence
211Exaone 4.0 32B0.0%estimated ± 12.3 pp, low confidence
212Gemini 2.5 Flash0.0%estimated ± 12.3 pp, low confidence
213Granite-4.0-350M0.0%estimated ± 12.3 pp, low confidence
214Granite-4.0-H-1B0.0%estimated ± 12.3 pp, low confidence
215Granite-4.0-H-350M0.0%estimated ± 12.3 pp, low confidence
216LFM2.5-VL-1.6B-Extract0.0%estimated ± 12.3 pp, low confidence
217Llama 3.1 405B0.0%estimated ± 12.3 pp, low confidence
218Mistral Medium 30.0%estimated ± 12.3 pp, low confidence
219Nemotron Ultra 253B0.0%estimated ± 12.3 pp, low confidence
220Nova Pro0.0%estimated ± 12.3 pp, low confidence
221Phi-40.0%estimated ± 12.3 pp, low confidence
222Qwen3-Omni-30B-A3B-Instruct0.0%estimated ± 12.3 pp, low confidence
223Qwen3-Omni-30B-A3B-Thinking0.0%estimated ± 12.3 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General