benchgap
Calibration

Terminal-Bench 4.0 → Terminal-Bench 2.1 (Vals AI)

Terminal-Bench 2.1 (Vals AI) is estimated from Terminal-Bench 4.0 with a Hill curve fitted on 5 models measured on both: y = 0.7610 + (1.2000 − 0.7610)·x^6.00 / (0.70056^6.00 + x^6.00), R² = 0.70, cross-validated error 4.6 pp. It is used for 12 estimates.

Estimated modelTerminal-Bench 4.0Terminal-Bench 2.1 (Vals AI)Source
Claude Haiku 5.539.2%77.4%estimated ± 4.6 pp, medium confidence
Claude Mythos 5.160.9%89.3%estimated ± 4.6 pp, low confidence
Claude Opus 5.566.4%94.5%estimated ± 4.6 pp, low confidence
Claude Sonnet 5.570.6%98.6%estimated ± 4.6 pp, low confidence
Gemini 4 Argon57.4%86.3%estimated ± 4.6 pp, medium confidence
Ling 3.1 Flash40.4%77.7%estimated ± 4.6 pp, medium confidence
MiMo-V2.6-Flash28.8%76.3%estimated ± 4.6 pp, medium confidence
MiMo-V2.6-Pro34.9%76.8%estimated ± 4.6 pp, medium confidence
Pareto 26.10 Preview50.8%81.7%estimated ± 4.6 pp, medium confidence
Pareto 26.951.0%81.8%estimated ± 4.6 pp, medium confidence
Step 5 Preview33.3%76.6%estimated ± 4.6 pp, medium confidence
SWE-227.3%76.2%estimated ± 4.6 pp, medium confidence