Calibration
Terminal-Bench 2.1 → Terminal-Bench 4.0
Terminal-Bench 4.0 is estimated from Terminal-Bench 2.1 with a inverse Michaelis–Menten curve fitted on 15 models measured on both: y = 6.64673·(x − 0.7648) / (0.7648 + 2.0000 − x), R² = 0.80, cross-validated error 11.6 pp. It is used for 4 estimates.
| Estimated model | Terminal-Bench 2.1 | Terminal-Bench 4.0 | Source |
|---|---|---|---|
| Claude Fable 5.1 (high with fallback) | 89.9% | 47.8% | estimated ± 11.6 pp, low confidence |
| GPT-6 Astra (high) | 89.9% | 47.8% | estimated ± 11.6 pp, low confidence |
| GPT-6 Astra (medium) | 89.5% | 46.3% | estimated ± 11.6 pp, low confidence |
| GPT-5.6 Sol (xhigh) | 89.5% | 46.3% | estimated ± 11.6 pp, low confidence |