benchgap
Agentic · terminal

Terminal-Bench-Science 0.1 leaderboard

As of 2026-10-07, the highest measured score on Terminal-Bench-Science 0.1 is 64.6% by GPT-6 Astra. 19 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra64.6%measured
2Claude Sonnet 5.559.9%measured
3Claude Opus 5.558.7%measured
4Gemini 4 Argon58.4%estimated ± 4.1 pp, medium confidence
5GPT-6.1 Sol57.0%measured
6Claude Fable 5.152.6%measured
7GPT-6 Sol39.5%estimated ± 4.1 pp, low confidence
8GLM-5.334.9%estimated ± 4.1 pp, low confidence
9MiMo-V2.6-Pro17.6%estimated ± 4.1 pp, low confidence
10Muse Spark 1.314.4%estimated ± 4.1 pp, low confidence
11Step 5 Preview14.4%estimated ± 4.1 pp, low confidence
12Claude Haiku 5.513.4%estimated ± 4.1 pp, low confidence
13GLM-5.3-Flash13.4%estimated ± 4.1 pp, low confidence
14DeepSeek V4.1 Flash4.7%estimated ± 4.1 pp, low confidence
15Mistral Large 44.7%estimated ± 4.1 pp, low confidence
16Grok 4.73.8%estimated ± 4.1 pp, low confidence
17Gemini 3.8 Flash0.8%estimated ± 4.1 pp, low confidence
18GPT-6 Luna0.1%estimated ± 4.1 pp, low confidence
19Kimi K30.1%estimated ± 4.1 pp, low confidence
20Qwen3.8-27B0.0%estimated ± 4.1 pp, low confidence
21Inkling0.0%estimated ± 4.1 pp, low confidence
22MiniMax M30.0%estimated ± 4.1 pp, low confidence
23Muse Glimmer 30B0.0%estimated ± 4.1 pp, low confidence
24Nemotron 3 Ultra0.0%estimated ± 4.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General