benchgap
Agentic · terminal

Terminal-Bench 2.1 leaderboard

As of 2026-10-07, the highest measured score on Terminal-Bench 2.1 is 91.4% by Claude Fable 5.1 (max with fallback). 17 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Claude Fable 5.1 (max with fallback)91.4%measured
2Claude Fable 5.1 (xhigh with fallback)91.0%measured
3Claude Fable 5.1 (high with fallback)89.9%measured
4GPT-6 Astra (high)89.9%measured
5GPT-6 Astra (medium)89.5%measured
6GPT-5.6 Sol (xhigh)89.5%measured
7Claude Opus 5.5 (medium with fallback)89.4%estimated ± 7.0 pp, medium confidence
8Claude Sonnet 5.5 (high with fallback)89.1%estimated ± 7.0 pp, medium confidence
9Qwen3.8 Max (0902)88.8%measured
10GPT-6 Astra (max)88.4%measured
11Claude Sonnet 5.5 (max with fallback)88.3%estimated ± 4.0 pp, medium confidence
12Claude Opus 5.5 (max with fallback)88.2%estimated ± 4.0 pp, medium confidence
13Claude Opus 5.5 (xhigh with fallback)88.2%estimated ± 4.0 pp, medium confidence
14GPT-6 Astra (xhigh)88.2%estimated ± 4.0 pp, medium confidence
15Gemini 4 Argon (high)88.2%estimated ± 4.0 pp, high confidence
16Claude Sonnet 5.5 (xhigh with fallback)88.2%estimated ± 4.0 pp, high confidence
17Claude Opus 5.5 (high with fallback)88.2%estimated ± 4.0 pp, high confidence
18GPT-6.1 Sol (max)88.2%estimated ± 4.0 pp, high confidence
19GPT-6 Sol (max)87.8%estimated ± 4.0 pp, high confidence
20Gemini 3.8 Flash (high)87.6%measured
21MiMo-V2.6-Pro87.4%estimated ± 4.0 pp, high confidence
22Step 5 Preview87.4%estimated ± 4.0 pp, high confidence
23DeepSeek V4.1 Flash (max)86.9%estimated ± 4.0 pp, high confidence
24Mistral Large 4 Preview86.9%estimated ± 4.0 pp, high confidence
25Grok 4.7 (xhigh)86.8%estimated ± 4.0 pp, high confidence
26Kimi K3 (max)85.0%measured
27GPT-6 Luna (max)84.3%estimated ± 4.0 pp, high confidence
28Muse Spark 1.3 (max)84.3%measured
29GLM-5.3-Flash84.3%measured
30GLM-5.3 (max)83.9%measured
31Qwen3.8 27B (xhigh)79.8%measured
32K2 Horizon 375B A23B71.9%measured
33MiniMax-M365.2%measured
34Inkling (xhigh)55.1%measured
35Nemotron 3 Ultra53.9%measured
36Muse Glimmer (high)51.7%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents