benchgap
Agentic · tools

APEX-Agents-AA leaderboard

As of 2026-10-07, the highest measured score on APEX-Agents-AA is 47.1% by Gemini 3.5 Flash. 189 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.548.1%estimated ± 9.1 pp, low confidence
2Gemini 3.5 Flash47.1%measured
3Claude Fable 545.0%estimated ± 3.5 pp, low confidence
4Qwen3.8 Max Preview42.4%measured
5MiMo-V2.6-Pro41.7%estimated ± 3.6 pp, low confidence
6Hy4 preview41.7%estimated ± 3.6 pp, low confidence
7MiMo-V2.6-Flash41.5%estimated ± 3.6 pp, low confidence
8Kimi K341.3%measured
9GPT-5.5 Pro41.2%estimated ± 4.6 pp, high confidence
10GPT-5.4 Pro39.9%estimated ± 4.6 pp, high confidence
11GPT-6.1 Sol39.6%estimated ± 9.1 pp, medium confidence
12Ling 3.1 Flash39.5%estimated ± 6.0 pp, low confidence
13GLM-5.3-Flash39.4%estimated ± 3.2 pp, medium confidence
14DeepSeek V4.1 Flash39.2%estimated ± 3.2 pp, medium confidence
15GPT-6 Astra38.9%estimated ± 2.9 pp, low confidence
16GPT-5.6 Terra38.9%measured
17Claude Opus 538.9%estimated ± 2.9 pp, low confidence
18Gemini 4 Argon38.9%estimated ± 2.9 pp, low confidence
19Muse Spark 1.338.8%estimated ± 2.9 pp, low confidence
20GPT-5.6 Sol38.7%estimated ± 2.9 pp, low confidence
21GPT-6 Sol38.7%estimated ± 2.9 pp, low confidence
22Fugu Cyber38.7%estimated ± 6.0 pp, low confidence
23Gemini 3.8 Flash38.7%estimated ± 2.9 pp, low confidence
24Claude Opus 5.538.4%estimated ± 2.9 pp, medium confidence
25Gemini 3.7 Flash38.4%estimated ± 2.9 pp, medium confidence
26GLM-5.338.4%estimated ± 3.2 pp, medium confidence
27Atria Dawn Preview38.3%estimated ± 3.6 pp, medium confidence
28Claude Fable 5.138.2%estimated ± 2.9 pp, medium confidence
29Gemini 3.8 Flash Cyber38.1%estimated ± 6.0 pp, low confidence
30Claude Sonnet 538.0%estimated ± 3.5 pp, low confidence
31Step 5 Preview38.0%measured
32Claude Mythos 537.8%estimated ± 4.6 pp, high confidence
33GPT-5.537.7%measured
34Grok 4.636.9%estimated ± 3.5 pp, low confidence
35Claude Opus 4.836.5%estimated ± 2.9 pp, medium confidence
36Muse Spark 1.236.4%estimated ± 9.1 pp, medium confidence
37Qwen3.8-Flash-Next36.3%estimated ± 2.9 pp, medium confidence
38Qwen3.8 Max36.3%estimated ± 2.9 pp, medium confidence
39Claude Opus 4.7 (Adaptive)36.1%estimated ± 2.9 pp, medium confidence
40Gemini 3.5 Flash Cyber35.9%estimated ± 6.0 pp, low confidence
41GPT-5.6 Luna35.8%measured
42Claude Mythos Preview35.8%estimated ± 6.0 pp, low confidence
43Ornith-1.5-397B35.6%estimated ± 4.6 pp, high confidence
44Muse Spark 1.135.2%estimated ± 2.9 pp, medium confidence
45GPT-6 Luna35.1%estimated ± 9.1 pp, medium confidence
46Claude Opus 4.735.1%estimated ± 2.9 pp, medium confidence
47Qwen3.7 Max34.8%estimated ± 5.7 pp, medium confidence
48Mistral Large 434.6%estimated ± 9.1 pp, medium confidence
49Kimi K2.7 Code34.6%estimated ± 5.7 pp, medium confidence
50Muse Glimmer 30B34.3%estimated ± 5.7 pp, medium confidence
51Claude Opus 4.633.8%estimated ± 3.5 pp, low confidence
52GLM-5.233.7%measured
53Grok 4.733.7%estimated ± 3.2 pp, low confidence
54Grok 4.533.5%estimated ± 9.1 pp, medium confidence
55GPT-5.433.3%measured
56Claude Opus 4.6 (Adaptive)33.0%measured
57Claude Sonnet 4.632.4%estimated ± 2.9 pp, medium confidence
58Gemini 3.1 Pro32.0%measured
59GPT-5.231.9%estimated ± 3.6 pp, medium confidence
60MiMo-V2.531.8%estimated ± 10.0 pp, low confidence
61GPT-5.3 Codex31.6%estimated ± 3.6 pp, medium confidence
62Qwen3.8-27B31.5%estimated ± 3.6 pp, medium confidence
63Apodex 1.131.2%measured
64Apodex 1.1 Mini31.2%measured
65Claude Opus 4.530.9%estimated ± 3.6 pp, medium confidence
66dots3-note Preview30.8%estimated ± 4.6 pp, high confidence
67Gemini 3.6 Flash30.2%estimated ± 9.1 pp, medium confidence
68MiMo-V2-Pro29.2%estimated ± 10.0 pp, low confidence
69Kimi K2.628.5%measured
70Claude Sonnet 4.528.4%estimated ± 3.6 pp, medium confidence
71Qwen3.6-35B-A3B28.3%estimated ± 5.7 pp, medium confidence
72GPT-5.4 mini28.2%measured
73MiniMax M328.2%estimated ± 2.9 pp, medium confidence
74Hy3 Preview27.9%estimated ± 9.1 pp, medium confidence
75GPT-5.1-Codex27.4%estimated ± 3.6 pp, medium confidence
76GPT-5.2-Codex27.3%estimated ± 3.6 pp, medium confidence
77Quasar 438B26.9%estimated ± 9.1 pp, medium confidence
78Ling 3.0 Flash VL26.2%estimated ± 9.1 pp, medium confidence
79Grok 4.126.2%estimated ± 10.0 pp, low confidence
80Solar Open 226.1%estimated ± 5.7 pp, medium confidence
81GLM-5-Turbo25.3%estimated ± 11.4 pp, low confidence
82GLM-5V-Turbo25.3%estimated ± 11.4 pp, low confidence
83Grok 4.1 Fast (Reasoning)25.3%estimated ± 11.4 pp, low confidence
84Qwen 3.6 Max (preview)25.3%estimated ± 11.4 pp, low confidence
85MiMo-V2-Omni25.3%estimated ± 11.4 pp, low confidence
86Claude Opus 4.5 Thinking25.3%estimated ± 11.4 pp, low confidence
87LFM2.5-8B-A1B25.2%estimated ± 11.4 pp, low confidence
88GPT-5.4 nano24.9%measured
89Claude 4.1 Opus24.6%estimated ± 3.6 pp, medium confidence
90GPT-5 (medium)24.5%estimated ± 11.4 pp, low confidence
91DeepSeek V4 Pro 081324.3%measured
92Inkling-Small23.2%estimated ± 4.6 pp, high confidence
93Beam23.2%estimated ± 4.6 pp, high confidence
94Hy323.0%estimated ± 9.1 pp, medium confidence
95Inkling22.9%estimated ± 4.6 pp, high confidence
96Qwen3.7 Plus22.4%measured
97Qwen3.5 Plus22.0%estimated ± 3.6 pp, medium confidence
98Claude 4 Sonnet21.9%estimated ± 3.6 pp, medium confidence
99Qwen3.6 Plus21.4%estimated ± 5.7 pp, medium confidence
100Agents-A121.1%estimated ± 4.6 pp, high confidence
101Qwen3.6-27B20.5%estimated ± 9.1 pp, medium confidence
102Gemini 3.5 Flash-Lite20.4%estimated ± 9.1 pp, medium confidence
103LLaDA2.2-flash20.4%estimated ± 5.7 pp, medium confidence
104Claude Haiku 4.519.9%estimated ± 3.6 pp, medium confidence
105A.X K219.7%estimated ± 9.1 pp, medium confidence
106Ling 3.0 Flash FP819.2%estimated ± 9.1 pp, medium confidence
107DeepSeek V4 Flash 073118.8%estimated ± 4.6 pp, high confidence
108Grok Build 0.118.5%estimated ± 11.4 pp, low confidence
109Ling 3.0 Flash17.9%estimated ± 4.6 pp, high confidence
110Grok 4.317.0%measured
111Gemini 3 Flash15.5%estimated ± 3.6 pp, medium confidence
112Gemini 3 Pro15.5%estimated ± 3.6 pp, medium confidence
113GPT-5.115.4%estimated ± 9.1 pp, medium confidence
114Step 3.7 Flash14.8%measured
115GLM-5.114.5%estimated ± 4.6 pp, high confidence
116GLM-514.5%measured
117Ornith-1.5-35B-A3B14.2%estimated ± 4.6 pp, high confidence
118Muse Spark14.0%estimated ± 6.0 pp, low confidence
119Agents-A1-4B13.7%estimated ± 4.6 pp, high confidence
120Mistral Medium 3.5 128B13.2%estimated ± 9.1 pp, medium confidence
121GPT-5 (high)12.3%estimated ± 3.6 pp, low confidence
122Qwen3.5-122B-A10B11.9%estimated ± 4.6 pp, high confidence
123Kimi K2.511.5%measured
124Kimi K2.5 (Reasoning)11.5%measured
125MiniCPM5-2B11.5%estimated ± 9.1 pp, medium confidence
126Qwen3.5 397B10.9%estimated ± 4.6 pp, high confidence
127Gemini 3.1 Flash-Lite10.6%estimated ± 11.4 pp, low confidence
128MiniMax M2.710.6%measured
129Grok 4.2010.5%estimated ± 11.4 pp, low confidence
130Qwen3.5-27B10.5%estimated ± 4.6 pp, high confidence
131Qwen3.5-35B-A3B10.5%estimated ± 4.6 pp, high confidence
132MiMo-V2-Flash9.0%estimated ± 9.1 pp, medium confidence
133Ornith-1.5-9B8.8%estimated ± 4.6 pp, medium confidence
134Gemma 4 31B8.6%estimated ± 9.1 pp, medium confidence
135GLM-4.77.6%estimated ± 4.6 pp, medium confidence
136Granite 4.2 30B7.2%estimated ± 9.1 pp, medium confidence
137Solar Pro 47.1%estimated ± 4.6 pp, medium confidence
138LongCat-Flash-Lite-Sparse7.0%estimated ± 4.6 pp, medium confidence
139Gemma 4 26B A4B6.9%estimated ± 9.1 pp, medium confidence
140Ling 3.0 Tiny6.8%estimated ± 9.1 pp, medium confidence
141Nemotron 3 Ultra6.5%estimated ± 4.6 pp, medium confidence
142Nemotron 3.5 Lightning 30B A3B NVFP46.0%estimated ± 4.6 pp, medium confidence
143Command A+5.3%estimated ± 9.1 pp, medium confidence
144Qwen3.5 397B (Reasoning)4.9%estimated ± 11.4 pp, low confidence
145Celeris-14.7%estimated ± 9.1 pp, medium confidence
146DeepSeek V34.7%estimated ± 9.1 pp, medium confidence
147DeepSeek V3 03244.7%estimated ± 9.1 pp, medium confidence
148Gemini 2.5 Pro4.7%estimated ± 9.1 pp, medium confidence
149Gemma 3 27B4.7%estimated ± 9.1 pp, medium confidence
150Gemma 4 12B4.7%estimated ± 9.1 pp, medium confidence
151Gemma 4 E2B4.7%estimated ± 9.1 pp, medium confidence
152Gemma 4 E4B4.7%estimated ± 9.1 pp, medium confidence
153GPT-4.1 mini4.7%estimated ± 9.1 pp, medium confidence
154GPT-4.1 nano4.7%estimated ± 9.1 pp, medium confidence
155GPT-4o4.7%estimated ± 9.1 pp, medium confidence
156GPT-4o mini4.7%estimated ± 9.1 pp, medium confidence
157Granite 4.2 3B4.7%estimated ± 9.1 pp, medium confidence
158Granite 4.2 8B4.7%estimated ± 9.1 pp, medium confidence
159K-Exaone4.7%estimated ± 9.1 pp, medium confidence
160LFM2.5-2.6B4.7%estimated ± 9.1 pp, medium confidence
161Ling 2.6 Flash4.7%estimated ± 9.1 pp, medium confidence
162Llama 4 Maverick4.7%estimated ± 9.1 pp, medium confidence
163Llama 4 Scout4.7%estimated ± 9.1 pp, medium confidence
164Mercury 2.54.7%estimated ± 9.1 pp, medium confidence
165Mistral Large 34.7%estimated ± 9.1 pp, medium confidence
166Mistral Small 44.7%estimated ± 9.1 pp, medium confidence
167Mistral Small 4 (Reasoning)4.7%estimated ± 9.1 pp, medium confidence
168Nemotron 3 Nano 30B4.7%estimated ± 9.1 pp, medium confidence
169Nemotron 3 Nano Omni 30B A3B4.7%estimated ± 9.1 pp, medium confidence
170North Mini Code4.7%estimated ± 9.1 pp, medium confidence
171Solar Pro 34.7%estimated ± 9.1 pp, medium confidence
172Trinity-Large-Preview4.7%estimated ± 9.1 pp, medium confidence
173Trinity-Large-Thinking4.7%estimated ± 9.1 pp, medium confidence
174Ultravox v0.6 Llama 3.3 70B4.7%estimated ± 9.1 pp, medium confidence
175GPT-OSS 120B3.1%measured
176MiMo-V2.5-Pro2.4%measured
177GPT-5.1-Codex-Max2.3%estimated ± 11.4 pp, low confidence
178Nemotron 3 Super 100B1.8%measured
179o31.7%estimated ± 11.4 pp, low confidence
180Ternary Bonsai 2 27B1.6%estimated ± 11.4 pp, low confidence
181DeepSeek V3.21.6%estimated ± 11.4 pp, low confidence
182Claude 3 Haiku1.6%estimated ± 11.4 pp, low confidence
183Claude 4.1 Opus Thinking1.6%estimated ± 11.4 pp, low confidence
184DeepSeek-R11.6%estimated ± 11.4 pp, low confidence
185DeepSeek V3.11.6%estimated ± 11.4 pp, low confidence
186DeepSeek V3.1 (Reasoning)1.6%estimated ± 11.4 pp, low confidence
187Exaone 4.0 1.2B1.6%estimated ± 11.4 pp, low confidence
188Exaone 4.0 32B1.6%estimated ± 11.4 pp, low confidence
189Gemini 2.5 Flash1.6%estimated ± 11.4 pp, low confidence
190GLM-4.5-Air1.6%estimated ± 11.4 pp, low confidence
191GLM-4.61.6%estimated ± 11.4 pp, low confidence
192GPT-4.11.6%estimated ± 11.4 pp, low confidence
193Granite-4.0-350M1.6%estimated ± 11.4 pp, low confidence
194Granite-4.0-H-1B1.6%estimated ± 11.4 pp, low confidence
195Granite-4.0-H-350M1.6%estimated ± 11.4 pp, low confidence
196Grok 41.6%estimated ± 11.4 pp, low confidence
197Grok 4.1 Fast1.6%estimated ± 11.4 pp, low confidence
198Grok 4 Fast (Reasoning)1.6%estimated ± 11.4 pp, low confidence
199Grok Code Fast 11.6%estimated ± 11.4 pp, low confidence
200Kimi K21.6%estimated ± 11.4 pp, low confidence
201LFM2.5-VL-1.6B-Extract1.6%estimated ± 11.4 pp, low confidence
202LLaDA2.2-mini1.6%estimated ± 11.4 pp, low confidence
203Llama 3.1 405B1.6%estimated ± 11.4 pp, low confidence
204Mistral Large 21.6%estimated ± 11.4 pp, low confidence
205Mistral Medium 31.6%estimated ± 11.4 pp, low confidence
206Nemotron Ultra 253B1.6%estimated ± 11.4 pp, low confidence
207Nova Pro1.6%estimated ± 11.4 pp, low confidence
208o11.6%estimated ± 11.4 pp, low confidence
209o3-mini1.6%estimated ± 11.4 pp, low confidence
210Phi-41.6%estimated ± 11.4 pp, low confidence
211Qwen3 Max1.6%estimated ± 11.4 pp, low confidence
212Qwen3-Omni-30B-A3B-Instruct1.6%estimated ± 11.4 pp, low confidence
213Qwen3-Omni-30B-A3B-Thinking1.6%estimated ± 11.4 pp, low confidence
214Sarvam 105B1.6%estimated ± 11.4 pp, low confidence
215Sarvam 30B1.6%estimated ± 11.4 pp, low confidence
216Solar Pro 21.6%estimated ± 11.4 pp, low confidence
217GPT-OSS 20B0.7%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General