benchgap
Agentic · terminal

Terminal-Bench Science 0.1 leaderboard

As of 2026-10-07, the highest measured score on Terminal-Bench Science 0.1 is 63.3% by GPT-6 Astra (max). 9 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1GPT-6 Astra (max)63.3%measured
2Claude Opus 5.5 (xhigh with fallback)61.9%measured
3Claude Opus 5.5 (max with fallback)59.0%measured
4GPT-6.1 Sol (max)58.1%measured
5GPT-6 Astra (xhigh)57.9%estimated ± 6.5 pp, medium confidence
6Gemini 4 Argon (high)53.4%estimated ± 6.5 pp, medium confidence
7Claude Sonnet 5.5 (max with fallback)53.3%measured
8Claude Sonnet 5.5 (xhigh with fallback)52.4%measured
9Claude Fable 5.1 (xhigh with fallback)49.4%estimated ± 6.5 pp, medium confidence
10Claude Opus 5.5 (high with fallback)49.0%measured
11Claude Fable 5.1 (max with fallback)43.3%measured
12Claude Opus 5.5 (medium with fallback)43.3%measured
13Claude Sonnet 5.5 (high with fallback)34.3%measured
14GPT-6 Sol (max)30.0%measured
15Qwen3.8 Max (0902)11.9%measured
16Muse Spark 1.3 (max)11.0%measured
17Gemini 3.8 Flash (high)10.0%measured
18GLM-5.3 (max)9.5%measured
19DeepSeek V4.1 Flash (max)9.0%measured
20GPT-6 Luna (max)8.6%measured
21MiMo-V2.6-Pro5.7%measured
22GLM-5.3-Flash4.8%measured
23Mistral Large 4 Preview4.1%estimated ± 6.5 pp, medium confidence
24Grok 4.7 (xhigh)3.8%estimated ± 6.5 pp, medium confidence
25Kimi K3 (max)2.6%estimated ± 6.5 pp, medium confidence
26K2 Horizon 375B A23B2.6%estimated ± 6.5 pp, medium confidence
27Nemotron 3 Ultra2.6%estimated ± 6.5 pp, low confidence
28Muse Glimmer (high)2.6%estimated ± 6.5 pp, low confidence
29Step 5 Preview2.4%measured
30Qwen3.8 27B (xhigh)0.5%measured
31MiniMax-M30.5%measured
32Inkling (xhigh)0.0%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents