LLM benchmark scores, with the gaps filled
benchgap is an LLM benchmark leaderboard that fills in the missing scores. Most models are only ever run on a handful of benchmarks, so benchgap calibrates benchmarks against each other on the models measured on both, then estimates each missing score with its cross-validated error and a confidence level. Measured and estimated scores are always marked apart.
As of 2026-10-07: 85 models, 24 benchmark versions, 396 measured and 207 estimated scores (58 high, 82 medium and 67 low confidence). Measured scores come from public leaderboards (artificial-analysis).
Terminal-Bench 4.0 leaderboard
As of 2026-10-07, the highest measured score on Terminal-Bench 4.0 is 63.6% by Claude Sonnet 5.5 (max with fallback). 6 more models have estimated scores, calibrated from the benchmarks they were measured on.
Measured scores: artificial-analysis.
All leaderboards
Agentic · terminal
Agentic · tools
- AA-AnalystAgent
- APEX-Agents
- AutomationBench
- EnterpriseOps-Gym
- Harvey LAB
- ITBench SRE
- τ³-Bench Banking
- τ²-Bench Telecom