benchgap
Agentic · tools

APEX-Agents leaderboard

As of 2026-10-07, the highest measured score on APEX-Agents is 57.5% by Grok 4.6. 125 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Grok 4.657.5%measured
2Claude Opus 5.552.3%estimated ± 10.0 pp, low confidence
3Claude Sonnet 5.551.5%estimated ± 10.0 pp, low confidence
4Claude Fable 5.149.1%estimated ± 10.0 pp, low confidence
5Grok 4.747.9%estimated ± 10.0 pp, low confidence
6MiMo-V2.6-Pro47.0%estimated ± 10.0 pp, low confidence
7Muse Spark 1.346.9%estimated ± 10.0 pp, low confidence
8Qwen3.8 Max Preview46.6%estimated ± 10.0 pp, low confidence
9GLM-5.346.0%estimated ± 10.0 pp, low confidence
10Qwen3.8-Flash-Next45.3%estimated ± 10.0 pp, low confidence
11Gemini 4 Argon45.1%estimated ± 10.0 pp, low confidence
12Ling 3.1 Flash45.0%estimated ± 10.0 pp, low confidence
13Claude Fable 544.7%estimated ± 10.0 pp, low confidence
14GPT-5.6 Sol44.7%estimated ± 10.0 pp, low confidence
15MiMo-V2.6-Flash44.6%estimated ± 10.0 pp, low confidence
16DeepSeek V4.1 Flash44.3%estimated ± 10.0 pp, low confidence
17GPT-6.1 Sol43.6%estimated ± 10.0 pp, low confidence
18GPT-6 Astra42.5%estimated ± 10.0 pp, low confidence
19GPT-6 Sol41.5%estimated ± 10.0 pp, low confidence
20Muse Spark 1.240.4%estimated ± 10.0 pp, low confidence
21Muse Spark 1.140.4%estimated ± 0.6 pp, low confidence
22Claude Sonnet 540.0%estimated ± 10.0 pp, low confidence
23GPT-5.6 Luna39.9%estimated ± 10.0 pp, low confidence
24GPT-5.6 Terra39.6%estimated ± 10.0 pp, low confidence
25GPT-6 Luna39.1%estimated ± 10.0 pp, low confidence
26Gemini 3.8 Flash39.0%estimated ± 10.0 pp, low confidence
27Mistral Large 438.6%estimated ± 10.0 pp, low confidence
28Qwen3.8-27B38.6%estimated ± 10.0 pp, low confidence
29Claude Opus 538.5%estimated ± 0.6 pp, low confidence
30Apodex 1.138.5%measured
31Step 5 Preview37.8%measured
32Kimi K337.6%measured
33Gemini 3.7 Flash37.5%estimated ± 10.0 pp, low confidence
34Grok 4.537.5%estimated ± 10.0 pp, low confidence
35Hy4 preview37.1%measured
36Gemini 3.5 Flash36.8%estimated ± 0.6 pp, medium confidence
37Claude Opus 4.835.6%estimated ± 0.6 pp, medium confidence
38Ornith-1.5-397B33.9%estimated ± 0.6 pp, medium confidence
39Gemini 3.6 Flash33.8%estimated ± 10.0 pp, low confidence
40Inkling-Small33.6%estimated ± 0.6 pp, medium confidence
41Beam32.8%estimated ± 0.6 pp, medium confidence
42Claude Opus 4.7 (Adaptive)31.7%estimated ± 0.6 pp, medium confidence
43GLM-5.231.3%estimated ± 0.6 pp, medium confidence
44Hy3 Preview31.3%estimated ± 10.0 pp, low confidence
45Qwen3.7 Max31.0%estimated ± 0.6 pp, medium confidence
46dots3-note Preview30.8%measured
47Kimi K2.7 Code30.7%estimated ± 0.6 pp, medium confidence
48Muse Glimmer 30B30.3%estimated ± 0.6 pp, medium confidence
49GPT-5.530.1%estimated ± 0.6 pp, medium confidence
50Quasar 438B30.1%estimated ± 10.0 pp, low confidence
51Ling 3.0 Flash VL29.3%estimated ± 10.0 pp, low confidence
52Nemotron 3 Ultra29.3%estimated ± 10.0 pp, low confidence
53MiniMax M329.2%estimated ± 0.6 pp, medium confidence
54Ling 3.0 Flash Fin29.2%measured
55Inkling29.1%estimated ± 0.6 pp, medium confidence
56DeepSeek V4 Pro 081328.7%estimated ± 0.6 pp, medium confidence
57Qwen3.7 Plus28.4%estimated ± 0.6 pp, medium confidence
58MiMo-V2.5-Pro27.8%estimated ± 10.0 pp, low confidence
59Apodex 1.1 Mini27.7%measured
60GLM-5.127.3%estimated ± 0.6 pp, medium confidence
61GPT-5.426.3%estimated ± 0.6 pp, medium confidence
62Grok 4.326.3%estimated ± 10.0 pp, low confidence
63Ornith-1.5-35B-A3B26.0%estimated ± 0.6 pp, medium confidence
64Hy325.6%estimated ± 10.0 pp, low confidence
65DeepSeek V4 Flash 073125.1%estimated ± 0.6 pp, medium confidence
66GLM-4.723.6%estimated ± 10.0 pp, low confidence
67Step 3.7 Flash23.6%estimated ± 10.0 pp, low confidence
68MiniMax M2.723.5%estimated ± 10.0 pp, low confidence
69Muse Spark23.0%estimated ± 10.0 pp, low confidence
70Qwen3.6-27B22.4%estimated ± 10.0 pp, low confidence
71Gemini 3.5 Flash-Lite22.4%estimated ± 10.0 pp, low confidence
72Ling 3.0 Flash22.3%estimated ± 0.6 pp, medium confidence
73A.X K221.5%estimated ± 10.0 pp, low confidence
74Ling 3.0 Flash FP820.8%estimated ± 10.0 pp, low confidence
75Qwen3.6-35B-A3B20.1%estimated ± 0.6 pp, medium confidence
76GPT-5 (high)19.6%estimated ± 10.0 pp, low confidence
77Solar Pro 418.7%measured
78Solar Open 216.6%measured
79Kimi K2.5 (Reasoning)16.4%estimated ± 10.0 pp, low confidence
80GPT-5.4 mini16.0%estimated ± 0.6 pp, low confidence
81GPT-5.115.7%estimated ± 10.0 pp, low confidence
82Qwen3.5-122B-A10B15.1%estimated ± 10.0 pp, low confidence
83GPT-5.4 nano14.7%estimated ± 0.6 pp, low confidence
84Kimi K2.614.6%estimated ± 0.6 pp, low confidence
85Gemini 3.1 Pro14.1%estimated ± 10.0 pp, low confidence
86Ornith-1.5-9B13.2%estimated ± 0.6 pp, low confidence
87Mistral Medium 3.5 128B12.8%estimated ± 10.0 pp, low confidence
88MiniCPM5-2B10.3%estimated ± 10.0 pp, low confidence
89Qwen3.6 Plus8.4%estimated ± 0.6 pp, low confidence
90Nemotron 3.5 Lightning 30B A3B NVFP47.0%estimated ± 10.0 pp, low confidence
91LLaDA2.2-flash6.8%estimated ± 0.6 pp, low confidence
92Qwen3.5 397B6.7%estimated ± 0.6 pp, low confidence
93MiMo-V2-Flash6.6%estimated ± 10.0 pp, low confidence
94LongCat-Flash-Lite-Sparse6.3%estimated ± 0.6 pp, low confidence
95Gemma 4 31B6.1%estimated ± 10.0 pp, low confidence
96GPT-OSS 120B5.6%estimated ± 10.0 pp, low confidence
97Granite 4.2 30B4.0%estimated ± 10.0 pp, low confidence
98Claude Opus 4.53.7%estimated ± 0.6 pp, low confidence
99Gemma 4 26B A4B3.5%estimated ± 10.0 pp, low confidence
100Ling 3.0 Tiny3.3%estimated ± 10.0 pp, low confidence
101Command A+0.9%estimated ± 10.0 pp, low confidence
102Celeris-10.0%estimated ± 10.0 pp, low confidence
103DeepSeek V30.0%estimated ± 10.0 pp, low confidence
104DeepSeek V3 03240.0%estimated ± 10.0 pp, low confidence
105Gemini 2.5 Pro0.0%estimated ± 10.0 pp, low confidence
106Gemma 3 27B0.0%estimated ± 10.0 pp, low confidence
107Gemma 4 12B0.0%estimated ± 10.0 pp, low confidence
108Gemma 4 E2B0.0%estimated ± 10.0 pp, low confidence
109Gemma 4 E4B0.0%estimated ± 10.0 pp, low confidence
110GLM-50.0%estimated ± 0.6 pp, low confidence
111GPT-4.1 mini0.0%estimated ± 10.0 pp, low confidence
112GPT-4.1 nano0.0%estimated ± 10.0 pp, low confidence
113GPT-4o0.0%estimated ± 10.0 pp, low confidence
114GPT-4o mini0.0%estimated ± 10.0 pp, low confidence
115GPT-OSS 20B0.0%estimated ± 10.0 pp, low confidence
116Granite 4.2 3B0.0%estimated ± 10.0 pp, low confidence
117Granite 4.2 8B0.0%estimated ± 10.0 pp, low confidence
118K-Exaone0.0%estimated ± 10.0 pp, low confidence
119Kimi K2.50.0%estimated ± 0.6 pp, low confidence
120LFM2.5-2.6B0.0%estimated ± 10.0 pp, low confidence
121Ling 2.6 Flash0.0%estimated ± 10.0 pp, low confidence
122Llama 4 Maverick0.0%estimated ± 10.0 pp, low confidence
123Llama 4 Scout0.0%estimated ± 10.0 pp, low confidence
124Mercury 2.50.0%estimated ± 10.0 pp, low confidence
125Mistral Large 30.0%estimated ± 10.0 pp, low confidence
126Mistral Small 40.0%estimated ± 10.0 pp, low confidence
127Mistral Small 4 (Reasoning)0.0%estimated ± 10.0 pp, low confidence
128Nemotron 3 Nano 30B0.0%estimated ± 10.0 pp, low confidence
129Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 10.0 pp, low confidence
130Nemotron 3 Super 100B0.0%estimated ± 10.0 pp, low confidence
131North Mini Code0.0%estimated ± 10.0 pp, low confidence
132Solar Pro 30.0%estimated ± 10.0 pp, low confidence
133Trinity-Large-Preview0.0%estimated ± 10.0 pp, low confidence
134Trinity-Large-Thinking0.0%estimated ± 10.0 pp, low confidence
135Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 10.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General