benchgap
Agentic · terminal

AA Terminal-Bench 2.1 leaderboard

As of 2026-10-07, the highest measured score on AA Terminal-Bench 2.1 is 91.4% by Claude Fable 5.1. 92 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.191.4%measured
2Claude Opus 588.8%estimated ± 4.4 pp, high confidence
3GPT-6 Astra88.4%measured
4Gemini 3.8 Flash87.6%measured
5Grok 4.687.6%estimated ± 4.4 pp, high confidence
6GPT-5.587.1%estimated ± 4.4 pp, high confidence
7Claude Sonnet 5.586.8%estimated ± 2.6 pp, medium confidence
8Claude Opus 5.586.8%estimated ± 2.6 pp, medium confidence
9Gemini 4 Argon86.8%estimated ± 2.6 pp, high confidence
10GPT-6.1 Sol86.8%estimated ± 2.6 pp, high confidence
11GPT-6 Sol86.8%estimated ± 2.6 pp, high confidence
12MiMo-V2.6-Pro86.7%estimated ± 2.6 pp, high confidence
13Step 5 Preview86.6%estimated ± 2.6 pp, high confidence
14Claude Haiku 5.586.6%estimated ± 2.6 pp, high confidence
15DeepSeek V4.1 Flash86.5%estimated ± 2.6 pp, high confidence
16Mistral Large 486.5%estimated ± 2.6 pp, high confidence
17Grok 4.786.4%estimated ± 2.6 pp, high confidence
18Gemini 3.6 Flash86.2%estimated ± 4.4 pp, high confidence
19GPT-6 Luna85.1%estimated ± 2.6 pp, high confidence
20Gemini 3.1 Pro85.0%estimated ± 4.4 pp, high confidence
21Kimi K385.0%measured
22SWE-284.9%estimated ± 3.5 pp, medium confidence
23GPT-5.6 Sol84.9%estimated ± 3.5 pp, medium confidence
24Claude Mythos 584.9%estimated ± 3.5 pp, high confidence
25DeepSeek V4 Pro 081384.9%estimated ± 3.5 pp, high confidence
26MiMo-V2.6-Flash84.9%estimated ± 3.5 pp, high confidence
27GPT-5.6 Terra84.9%estimated ± 3.5 pp, high confidence
28Qwen3.8 Max84.9%estimated ± 3.5 pp, high confidence
29Ornith-1.5-397B84.9%estimated ± 3.5 pp, high confidence
30Gemini 3.7 Flash84.9%estimated ± 3.5 pp, high confidence
31Hy4 preview84.9%estimated ± 3.5 pp, high confidence
32GPT-5.6 Luna84.9%estimated ± 3.5 pp, high confidence
33Claude Fable 584.9%estimated ± 3.5 pp, high confidence
34Grok 4.584.8%estimated ± 3.5 pp, high confidence
35Muse Spark 1.284.8%estimated ± 3.5 pp, high confidence
36DeepSeek V4 Flash 073184.8%estimated ± 3.5 pp, high confidence
37Sakana Fugu-Ultra84.8%estimated ± 3.5 pp, high confidence
38Ember-184.8%estimated ± 3.5 pp, high confidence
39SWE-1.784.7%estimated ± 3.5 pp, high confidence
40GLM-5.284.7%estimated ± 3.5 pp, high confidence
41Claude Sonnet 584.6%estimated ± 3.5 pp, high confidence
42Sakana Fugu84.6%estimated ± 3.5 pp, high confidence
43Beam84.6%estimated ± 3.5 pp, high confidence
44Muse Spark 1.184.6%estimated ± 3.5 pp, high confidence
45GLM-5.3-Flash84.3%measured
46Muse Spark 1.384.3%measured
47Atria Dawn Preview84.3%estimated ± 3.5 pp, high confidence
48Ornith-1.0-397B84.1%estimated ± 3.5 pp, high confidence
49GLM-5.383.9%measured
50Claude Opus 4.783.8%estimated ± 4.4 pp, high confidence
51Gemini 3.5 Flash83.5%estimated ± 3.5 pp, high confidence
52Kimi K2.7 Code82.9%estimated ± 4.4 pp, high confidence
53dots3-note Preview82.9%estimated ± 3.5 pp, high confidence
54Claude Opus 4.882.5%estimated ± 3.5 pp, high confidence
55Qwen3.8-27B79.8%measured
56Qwen3.7 Max77.8%estimated ± 4.4 pp, high confidence
57MiMo-V2.577.5%estimated ± 4.4 pp, high confidence
58Seed 2.1 Pro77.2%estimated ± 3.5 pp, high confidence
59Apodex 1.176.7%estimated ± 3.5 pp, high confidence
60Laguna S 2.175.2%estimated ± 3.5 pp, high confidence
61Composer 2.574.6%estimated ± 4.4 pp, high confidence
62Claude Sonnet 4.673.0%estimated ± 4.4 pp, high confidence
63MiMo-V2.5-Pro73.0%estimated ± 4.4 pp, high confidence
64Quasar 438B72.8%estimated ± 3.5 pp, high confidence
65GLM-5.172.4%estimated ± 4.4 pp, high confidence
66GPT-5.4 mini68.7%estimated ± 4.4 pp, high confidence
67Ornith-1.5-35B-A3B68.3%estimated ± 3.5 pp, high confidence
68Seed 2.1 Turbo67.6%estimated ± 3.5 pp, high confidence
69Gemini 3 Flash67.2%estimated ± 4.4 pp, high confidence
70Kimi K2.666.6%estimated ± 4.4 pp, high confidence
71Qwen3.6 Plus65.8%estimated ± 4.4 pp, high confidence
72MiniMax M365.2%measured
73Qwen3.7 Plus64.9%estimated ± 4.4 pp, high confidence
74Pokee-Isaac 28B60.6%estimated ± 3.5 pp, high confidence
75Inkling-Small59.7%estimated ± 3.5 pp, high confidence
76Ornith-1.0-35B58.7%estimated ± 3.5 pp, high confidence
77MAI-Code-1.1-Flash56.5%estimated ± 3.5 pp, high confidence
78MiniMax M2.755.1%estimated ± 4.4 pp, high confidence
79Inkling55.1%measured
80Nemotron 3 Ultra53.9%measured
81Step 3.7 Flash53.4%estimated ± 3.5 pp, high confidence
82Ling 3.0 Flash52.6%estimated ± 3.5 pp, high confidence
83Solar Pro 452.6%estimated ± 3.5 pp, high confidence
84Gemini 3.5 Flash-Lite52.3%estimated ± 3.5 pp, high confidence
85Ternary Bonsai 2 27B52.2%estimated ± 3.5 pp, high confidence
86Ornith-1.5-9B52.1%estimated ± 3.5 pp, medium confidence
87K-EXAONE 2.052.1%estimated ± 3.5 pp, medium confidence
88Ornith-1.0-9B52.1%estimated ± 3.5 pp, medium confidence
89A.X K252.1%estimated ± 3.5 pp, medium confidence
90Granite 4.2 30B52.1%estimated ± 3.5 pp, medium confidence
91Granite 4.2 8B52.1%estimated ± 3.5 pp, medium confidence
92MiniCPM5-2B52.1%estimated ± 3.5 pp, medium confidence
93Nemotron 3.5 Lightning 30B A3B NVFP452.1%estimated ± 3.5 pp, medium confidence
94Muse Glimmer 30B51.7%measured
95Grok 4.2042.1%estimated ± 4.4 pp, medium confidence
96Claude Haiku 4.540.8%estimated ± 4.4 pp, medium confidence
97Grok 4.335.0%estimated ± 4.4 pp, medium confidence
98GPT-5.4 nano34.0%estimated ± 4.4 pp, medium confidence
99Mistral Medium 3.5 128B26.3%estimated ± 4.4 pp, medium confidence
100Mercury 2.514.8%estimated ± 4.4 pp, medium confidence
101Gemini 3.1 Flash-Lite14.0%estimated ± 4.4 pp, medium confidence
102Laguna M.114.0%estimated ± 4.4 pp, medium confidence
103Laguna XS.23.0%estimated ± 4.4 pp, medium confidence
104Command A+0.3%estimated ± 4.4 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General