benchgap
Agentic · terminal

AA Terminal-Bench 4.0 leaderboard

As of 2026-10-07, the highest measured score on AA Terminal-Bench 4.0 is 63.6% by Claude Sonnet 5.5. 84 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.563.6%measured
2Claude Opus 5.559.6%measured
3GPT-6 Astra59.1%measured
4Gemini 4 Argon57.1%measured
5Claude Mythos 5.156.5%estimated ± 4.7 pp, high confidence
6GPT-6.1 Sol56.1%measured
7Claude Fable 5.152.0%measured
8Pareto 26.947.1%estimated ± 4.7 pp, high confidence
9Pareto 26.10 Preview47.0%estimated ± 4.7 pp, high confidence
10GPT-6 Sol43.9%measured
11GLM-5.341.9%measured
12Claude Opus 541.5%estimated ± 14.0 pp, low confidence
13Ling 3.1 Flash37.1%estimated ± 4.7 pp, high confidence
14MiMo-V2.6-Pro34.8%measured
15Grok 4.634.4%estimated ± 14.0 pp, low confidence
16GPT-5.6 Sol33.8%estimated ± 8.9 pp, low confidence
17Muse Spark 1.333.3%measured
18Step 5 Preview33.3%measured
19Claude Haiku 5.532.8%measured
20GLM-5.3-Flash32.8%measured
21GPT-5.532.2%estimated ± 14.0 pp, low confidence
22Gemini 3.6 Flash29.3%estimated ± 14.0 pp, low confidence
23Claude Mythos 528.6%estimated ± 8.9 pp, medium confidence
24DeepSeek V4 Pro 081328.5%estimated ± 8.9 pp, medium confidence
25GPT-5.6 Terra27.8%estimated ± 8.9 pp, medium confidence
26DeepSeek V4.1 Flash26.8%measured
27Mistral Large 426.8%measured
28Qwen3.8 Max26.8%estimated ± 8.9 pp, medium confidence
29MiMo-V2.6-Flash26.2%estimated ± 4.7 pp, high confidence
30Ornith-1.5-397B26.1%estimated ± 8.9 pp, medium confidence
31Gemini 3.1 Pro25.9%estimated ± 14.0 pp, low confidence
32Grok 4.725.8%measured
33Gemini 3.7 Flash25.7%estimated ± 8.9 pp, medium confidence
34Hy4 preview25.2%estimated ± 8.9 pp, medium confidence
35SWE-224.8%estimated ± 4.7 pp, high confidence
36GPT-5.6 Luna24.3%estimated ± 8.9 pp, medium confidence
37Claude Fable 523.8%estimated ± 8.9 pp, medium confidence
38Claude Opus 4.723.3%estimated ± 14.0 pp, low confidence
39Grok 4.522.6%estimated ± 8.9 pp, medium confidence
40Muse Spark 1.222.1%estimated ± 8.9 pp, medium confidence
41DeepSeek V4 Flash 073121.8%estimated ± 8.9 pp, medium confidence
42Kimi K2.7 Code21.6%estimated ± 14.0 pp, low confidence
43Sakana Fugu-Ultra21.1%estimated ± 8.9 pp, medium confidence
44Ember-121.0%estimated ± 8.9 pp, medium confidence
45SWE-1.720.4%estimated ± 8.9 pp, medium confidence
46GLM-5.219.8%estimated ± 8.9 pp, medium confidence
47Gemini 3.8 Flash19.7%measured
48Claude Sonnet 519.0%estimated ± 8.9 pp, medium confidence
49Sakana Fugu18.8%estimated ± 8.9 pp, medium confidence
50Beam18.7%estimated ± 8.9 pp, medium confidence
51Muse Spark 1.118.5%estimated ± 8.9 pp, medium confidence
52Atria Dawn Preview16.5%estimated ± 8.9 pp, medium confidence
53Ornith-1.0-397B15.6%estimated ± 8.9 pp, medium confidence
54Qwen3.7 Max14.8%estimated ± 14.0 pp, low confidence
55MiMo-V2.514.5%estimated ± 14.0 pp, low confidence
56Gemini 3.5 Flash14.1%estimated ± 8.9 pp, medium confidence
57dots3-note Preview12.8%estimated ± 8.9 pp, medium confidence
58GPT-6 Luna12.6%measured
59Kimi K312.6%measured
60Claude Opus 4.812.2%estimated ± 8.9 pp, medium confidence
61Composer 2.511.9%estimated ± 14.0 pp, low confidence
62Claude Sonnet 4.610.6%estimated ± 14.0 pp, low confidence
63MiMo-V2.5-Pro10.6%estimated ± 14.0 pp, low confidence
64GLM-5.110.2%estimated ± 14.0 pp, low confidence
65Seed 2.1 Pro8.2%estimated ± 8.9 pp, medium confidence
66Apodex 1.18.0%estimated ± 8.9 pp, medium confidence
67GPT-5.4 mini7.7%estimated ± 14.0 pp, low confidence
68Laguna S 2.17.4%estimated ± 8.9 pp, medium confidence
69Gemini 3 Flash6.8%estimated ± 14.0 pp, low confidence
70Kimi K2.66.5%estimated ± 14.0 pp, low confidence
71Quasar 438B6.4%estimated ± 8.9 pp, medium confidence
72Qwen3.6 Plus6.0%estimated ± 14.0 pp, low confidence
73Qwen3.8-27B5.6%measured
74Qwen3.7 Plus5.6%estimated ± 14.0 pp, low confidence
75Ornith-1.5-35B-A3B4.8%estimated ± 8.9 pp, medium confidence
76Seed 2.1 Turbo4.6%estimated ± 8.9 pp, medium confidence
77MiniMax M32.0%measured
78Pokee-Isaac 28B2.0%estimated ± 8.9 pp, medium confidence
79Inkling-Small1.6%estimated ± 8.9 pp, medium confidence
80Ornith-1.0-35B1.1%estimated ± 8.9 pp, medium confidence
81Inkling1.0%measured
82MiniMax M2.70.9%estimated ± 14.0 pp, low confidence
83Muse Glimmer 30B0.5%measured
84Nemotron 3 Ultra0.5%measured
85A.X K20.0%estimated ± 8.9 pp, low confidence
86Claude Haiku 4.50.0%estimated ± 14.0 pp, low confidence
87Command A+0.0%estimated ± 14.0 pp, low confidence
88Gemini 3.1 Flash-Lite0.0%estimated ± 14.0 pp, low confidence
89Gemini 3.5 Flash-Lite0.0%estimated ± 8.9 pp, medium confidence
90GPT-5.4 nano0.0%estimated ± 14.0 pp, low confidence
91Granite 4.2 30B0.0%estimated ± 8.9 pp, low confidence
92Granite 4.2 8B0.0%estimated ± 8.9 pp, low confidence
93Grok 4.200.0%estimated ± 14.0 pp, low confidence
94Grok 4.30.0%estimated ± 14.0 pp, low confidence
95K-EXAONE 2.00.0%estimated ± 8.9 pp, low confidence
96Laguna M.10.0%estimated ± 14.0 pp, low confidence
97Laguna XS.20.0%estimated ± 14.0 pp, low confidence
98Ling 3.0 Flash0.0%estimated ± 8.9 pp, medium confidence
99MAI-Code-1.1-Flash0.0%estimated ± 8.9 pp, medium confidence
100Mercury 2.50.0%estimated ± 14.0 pp, low confidence
101MiniCPM5-2B0.0%estimated ± 8.9 pp, low confidence
102Mistral Medium 3.5 128B0.0%estimated ± 14.0 pp, low confidence
103Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 8.9 pp, low confidence
104Ornith-1.0-9B0.0%estimated ± 8.9 pp, low confidence
105Ornith-1.5-9B0.0%estimated ± 8.9 pp, low confidence
106Solar Pro 40.0%estimated ± 8.9 pp, medium confidence
107Step 3.7 Flash0.0%estimated ± 8.9 pp, medium confidence
108Ternary Bonsai 2 27B0.0%estimated ± 8.9 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General