benchgap
Calibration

Terminal-Bench 2.1 → Terminal-Bench 4.0

Terminal-Bench 4.0 is estimated from Terminal-Bench 2.1 with a inverse Michaelis–Menten curve fitted on 15 models measured on both: y = 6.64673·(x − 0.7648) / (0.7648 + 2.0000 − x), R² = 0.80, cross-validated error 11.6 pp. It is used for 4 estimates.

Estimated modelTerminal-Bench 2.1Terminal-Bench 4.0Source
Claude Fable 5.1 (high with fallback)89.9%47.8%estimated ± 11.6 pp, low confidence
GPT-6 Astra (high)89.9%47.8%estimated ± 11.6 pp, low confidence
GPT-6 Astra (medium)89.5%46.3%estimated ± 11.6 pp, low confidence
GPT-5.6 Sol (xhigh)89.5%46.3%estimated ± 11.6 pp, low confidence