benchgap
Calibration

AA Terminal-Bench 4.0 → Terminal-Bench 2.1

Terminal-Bench 2.1 is estimated from AA Terminal-Bench 4.0 with a Michaelis–Menten + offset curve fitted on 13 models measured on both: y = 0.4818 + 0.4344·x / (0.02665 + x), R² = 0.95, cross-validated error 3.9 pp. It is used for 11 estimates.

Estimated modelAA Terminal-Bench 4.0Terminal-Bench 2.1Source
Claude Fable 5.152.0%89.5%estimated ± 3.9 pp, medium confidence
Claude Haiku 5.532.8%88.4%estimated ± 3.9 pp, high confidence
Claude Opus 5.559.6%89.8%estimated ± 3.9 pp, medium confidence
Claude Sonnet 5.563.6%89.9%estimated ± 3.9 pp, medium confidence
Gemini 4 Argon57.1%89.7%estimated ± 3.9 pp, medium confidence
GPT-6.1 Sol56.1%89.6%estimated ± 3.9 pp, medium confidence
GPT-6 Astra59.1%89.7%estimated ± 3.9 pp, medium confidence
GPT-6 Luna12.6%84.0%estimated ± 3.9 pp, high confidence
GPT-6 Sol43.9%89.1%estimated ± 3.9 pp, medium confidence
Grok 4.725.8%87.6%estimated ± 3.9 pp, high confidence
Mistral Large 426.8%87.7%estimated ± 3.9 pp, high confidence