benchgap
Agentic · terminal

Terminal-Bench 3.0 leaderboard

As of 2026-10-07, the highest measured score on Terminal-Bench 3.0 is 42.7% by Claude Opus 5. 83 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 542.7%measured
2GPT-6 Astra38.5%estimated ± 6.9 pp, low confidence
3SWE-236.4%estimated ± 9.1 pp, low confidence
4GPT-5.6 Sol34.6%measured
5Claude Fable 5.134.5%estimated ± 6.9 pp, medium confidence
6Claude Fable 534.1%measured
7Gemini 3.8 Flash28.3%estimated ± 6.9 pp, medium confidence
8Kimi K327.6%estimated ± 6.9 pp, medium confidence
9MiMo-V2.6-Pro27.0%estimated ± 9.1 pp, low confidence
10Grok 4.626.5%measured
11Claude Mythos 523.0%estimated ± 9.1 pp, low confidence
12MiMo-V2.6-Flash22.2%estimated ± 9.1 pp, low confidence
13Claude Opus 4.821.1%measured
14GPT-5.521.0%estimated ± 6.9 pp, medium confidence
15GPT-5.6 Terra20.8%measured
16DeepSeek V4.1 Flash18.5%estimated ± 6.9 pp, medium confidence
17Step 5 Preview18.3%estimated ± 9.1 pp, low confidence
18Gemini 3.5 Flash18.2%estimated ± 6.9 pp, medium confidence
19Gemini 3.6 Flash17.7%estimated ± 6.9 pp, medium confidence
20Grok 4.717.2%estimated ± 6.9 pp, medium confidence
21Muse Spark 1.315.9%estimated ± 6.9 pp, medium confidence
22Grok 4.515.7%measured
23Sakana Fugu-Ultra15.2%estimated ± 9.1 pp, low confidence
24Ember-115.1%estimated ± 9.1 pp, low confidence
25GLM-5.315.0%estimated ± 6.9 pp, medium confidence
26Gemini 3.7 Flash14.9%measured
27SWE-1.714.6%estimated ± 9.1 pp, low confidence
28Claude Sonnet 514.6%measured
29GPT-5.6 Luna14.3%measured
30Gemini 3.1 Pro14.2%estimated ± 6.9 pp, medium confidence
31Sakana Fugu13.6%estimated ± 9.1 pp, low confidence
32Ornith-1.5-397B13.5%measured
33Beam13.5%estimated ± 9.1 pp, low confidence
34Muse Spark 1.213.1%estimated ± 6.9 pp, medium confidence
35Muse Spark 1.112.7%estimated ± 6.9 pp, medium confidence
36Atria Dawn Preview12.2%estimated ± 9.1 pp, low confidence
37Claude Opus 4.711.9%estimated ± 6.9 pp, medium confidence
38Ornith-1.0-397B11.7%estimated ± 9.1 pp, low confidence
39Qwen3.8 Max10.9%estimated ± 6.9 pp, low confidence
40DeepSeek V4 Flash 073110.6%estimated ± 6.9 pp, low confidence
41Kimi K2.7 Code10.6%estimated ± 6.9 pp, low confidence
42dots3-note Preview10.3%estimated ± 9.1 pp, low confidence
43Seed 2.1 Pro8.5%estimated ± 9.1 pp, low confidence
44Apodex 1.18.4%estimated ± 9.1 pp, low confidence
45Laguna S 2.18.2%estimated ± 9.1 pp, low confidence
46Quasar 438B7.8%estimated ± 9.1 pp, low confidence
47GLM-5.3-Flash7.4%estimated ± 6.9 pp, low confidence
48Seed 2.1 Turbo7.3%estimated ± 9.1 pp, low confidence
49Pokee-Isaac 28B6.5%estimated ± 9.1 pp, low confidence
50Ornith-1.0-35B6.3%estimated ± 9.1 pp, low confidence
51Qwen3.7 Max6.3%estimated ± 6.9 pp, low confidence
52MiMo-V2.56.1%estimated ± 6.9 pp, low confidence
53MAI-Code-1.1-Flash6.0%estimated ± 9.1 pp, low confidence
54Step 3.7 Flash5.2%estimated ± 9.1 pp, low confidence
55Ornith-1.5-35B-A3B5.1%measured
56Composer 2.54.9%estimated ± 6.9 pp, low confidence
57Qwen3.8-27B4.9%estimated ± 6.9 pp, low confidence
58Solar Pro 44.7%estimated ± 9.1 pp, low confidence
59GLM-5.24.6%measured
60Claude Sonnet 4.64.4%estimated ± 6.9 pp, low confidence
61MiMo-V2.5-Pro4.4%estimated ± 6.9 pp, low confidence
62GLM-5.14.2%estimated ± 6.9 pp, low confidence
63Ternary Bonsai 2 27B4.0%estimated ± 9.1 pp, low confidence
64Muse Glimmer 30B3.8%estimated ± 9.1 pp, low confidence
65Hy4 preview3.5%estimated ± 6.9 pp, low confidence
66Inkling-Small3.5%estimated ± 6.9 pp, low confidence
67DeepSeek V4 Pro 08133.3%estimated ± 6.9 pp, low confidence
68GPT-5.4 mini3.3%estimated ± 6.9 pp, low confidence
69Ornith-1.5-9B3.1%estimated ± 9.1 pp, low confidence
70Gemini 3 Flash3.1%estimated ± 6.9 pp, low confidence
71Kimi K2.63.0%estimated ± 6.9 pp, low confidence
72MiniMax M33.0%estimated ± 6.9 pp, low confidence
73Qwen3.6 Plus2.8%estimated ± 6.9 pp, low confidence
74K-EXAONE 2.02.8%estimated ± 9.1 pp, low confidence
75Ornith-1.0-9B2.7%estimated ± 9.1 pp, low confidence
76Qwen3.7 Plus2.7%estimated ± 6.9 pp, low confidence
77Nemotron 3 Ultra2.2%estimated ± 6.9 pp, low confidence
78A.X K22.0%estimated ± 9.1 pp, low confidence
79Gemini 3.5 Flash-Lite2.0%estimated ± 6.9 pp, low confidence
80Ling 3.0 Flash2.0%estimated ± 6.9 pp, low confidence
81MiniMax M2.71.7%estimated ± 6.9 pp, low confidence
82Granite 4.2 30B1.5%estimated ± 9.1 pp, low confidence
83Inkling1.5%estimated ± 6.9 pp, low confidence
84Nemotron 3.5 Lightning 30B A3B NVFP41.1%estimated ± 9.1 pp, low confidence
85Grok 4.200.9%estimated ± 6.9 pp, low confidence
86Granite 4.2 8B0.9%estimated ± 9.1 pp, low confidence
87Claude Haiku 4.50.9%estimated ± 6.9 pp, low confidence
88Grok 4.30.7%estimated ± 6.9 pp, low confidence
89GPT-5.4 nano0.7%estimated ± 6.9 pp, low confidence
90Mistral Medium 3.5 128B0.4%estimated ± 6.9 pp, low confidence
91MiniCPM5-2B0.3%estimated ± 9.1 pp, low confidence
92Mercury 2.50.2%estimated ± 6.9 pp, low confidence
93Gemini 3.1 Flash-Lite0.2%estimated ± 6.9 pp, low confidence
94Laguna M.10.2%estimated ± 6.9 pp, low confidence
95Laguna XS.20.0%estimated ± 6.9 pp, low confidence
96Command A+0.0%estimated ± 6.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General