benchgap
LLM benchmark leaderboard · gaps filled

LLM benchmark scores, with the gaps filled

benchgap is an LLM benchmark leaderboard that fills in the missing scores. Most models are only ever run on a handful of benchmarks, so benchgap calibrates benchmarks against each other on the models measured on both, then estimates each missing score with its cross-validated error and a confidence level. Measured and estimated scores are always marked apart.

As of 2026-10-07: 401 models, 147 benchmark versions, 4892 measured and 16443 estimated scores (2500 high, 5182 medium and 8761 low confidence). Measured scores come from public leaderboards (artificial-analysis, published, vals-ai).

Terminal-Bench 4.0 leaderboard

As of 2026-10-07, the highest measured score on Terminal-Bench 4.0 is 70.6% by Claude Sonnet 5.5. 13 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.570.6%measured
2Claude Opus 5.566.4%measured
3Claude Mythos 5.160.9%measured
4GPT-6.1 Sol59.5%estimated ± 5.1 pp, medium confidence
5GPT-6 Astra57.9%measured
6Gemini 4 Argon57.4%measured
7Claude Fable 5.155.8%measured
8Pareto 26.951.0%measured
9Pareto 26.10 Preview50.8%measured
10GPT-6 Sol47.4%estimated ± 5.1 pp, medium confidence
11GLM-5.345.5%estimated ± 5.1 pp, medium confidence
12Ling 3.1 Flash40.4%measured
13Claude Haiku 5.539.2%measured
14Grok 4.738.0%measured
15Muse Spark 1.337.0%estimated ± 5.1 pp, medium confidence
16GLM-5.3-Flash36.5%estimated ± 5.1 pp, medium confidence
17MiMo-V2.6-Pro34.9%measured
18Step 5 Preview33.3%measured
19DeepSeek V4.1 Flash31.2%measured
20Mistral Large 430.5%estimated ± 5.1 pp, medium confidence
21MiMo-V2.6-Flash28.8%measured
22SWE-227.3%measured
23Gemini 3.8 Flash19.1%measured
24GPT-6 Luna16.5%estimated ± 5.1 pp, low confidence
25Kimi K316.5%estimated ± 5.1 pp, low confidence
26Qwen3.8-27B9.5%estimated ± 5.1 pp, low confidence
27MiniMax M36.0%estimated ± 5.1 pp, low confidence
28Inkling5.0%estimated ± 5.1 pp, low confidence
29Muse Glimmer 30B4.5%estimated ± 5.1 pp, low confidence
30Nemotron 3 Ultra4.5%estimated ± 5.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General