benchgap
LLM benchmark leaderboard · gaps filled

LLM benchmark scores, with the gaps filled

benchgap is an LLM benchmark leaderboard that fills in the missing scores. Most models are only ever run on a handful of benchmarks, so benchgap calibrates benchmarks against each other on the models measured on both, then estimates each missing score with its cross-validated error and a confidence level. Measured and estimated scores are always marked apart.

As of 2026-10-07: 85 models, 24 benchmark versions, 396 measured and 207 estimated scores (58 high, 82 medium and 67 low confidence). Measured scores come from public leaderboards (artificial-analysis).

Terminal-Bench 4.0 leaderboard

As of 2026-10-07, the highest measured score on Terminal-Bench 4.0 is 63.6% by Claude Sonnet 5.5 (max with fallback). 6 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Claude Sonnet 5.5 (max with fallback)63.6%measured
2Claude Opus 5.5 (max with fallback)59.6%measured
3Claude Opus 5.5 (xhigh with fallback)59.6%measured
4GPT-6 Astra (xhigh)59.6%measured
5GPT-6 Astra (max)59.1%measured
6Gemini 4 Argon (high)57.1%measured
7Claude Sonnet 5.5 (xhigh with fallback)57.1%measured
8Claude Opus 5.5 (high with fallback)56.6%measured
9GPT-6.1 Sol (max)56.1%measured
10Claude Fable 5.1 (xhigh with fallback)55.1%measured
11Claude Opus 5.5 (medium with fallback)53.6%estimated ± 8.2 pp, medium confidence
12Claude Fable 5.1 (max with fallback)52.0%measured
13Claude Sonnet 5.5 (high with fallback)50.0%estimated ± 8.2 pp, medium confidence
14Claude Fable 5.1 (high with fallback)47.8%estimated ± 11.6 pp, low confidence
15GPT-6 Astra (high)47.8%estimated ± 11.6 pp, low confidence
16GPT-6 Astra (medium)46.3%estimated ± 11.6 pp, low confidence
17GPT-5.6 Sol (xhigh)46.3%estimated ± 11.6 pp, low confidence
18GPT-6 Sol (max)43.9%measured
19GLM-5.3 (max)41.9%measured
20Qwen3.8 Max (0902)38.9%measured
21MiMo-V2.6-Pro34.8%measured
22Muse Spark 1.3 (max)33.3%measured
23Step 5 Preview33.3%measured
24GLM-5.3-Flash32.8%measured
25DeepSeek V4.1 Flash (max)26.8%measured
26Mistral Large 4 Preview26.8%measured
27Grok 4.7 (xhigh)25.8%measured
28Gemini 3.8 Flash (high)19.7%measured
29Kimi K3 (max)12.6%measured
30GPT-6 Luna (max)12.6%measured
31Qwen3.8 27B (xhigh)5.6%measured
32MiniMax-M32.0%measured
33K2 Horizon 375B A23B1.5%measured
34Inkling (xhigh)1.0%measured
35Nemotron 3 Ultra0.5%measured
36Muse Glimmer (high)0.5%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents