benchgap
Calibration

Terminal-Bench 2.1 (Vals AI) → Terminal-Bench 2.1

Terminal-Bench 2.1 is estimated from Terminal-Bench 2.1 (Vals AI) with a Hill curve fitted on 29 models measured on both: y = 0.0000 + (0.8870 − 0.0000)·x^6.00 / (0.43383^6.00 + x^6.00), R² = 0.67, cross-validated error 6.9 pp. It is used for 27 estimates.

Estimated modelTerminal-Bench 2.1 (Vals AI)Terminal-Bench 2.1Source
Claude Haiku 4.543.8%45.6%estimated ± 6.9 pp, low confidence
Claude Opus 4.768.5%83.3%estimated ± 6.9 pp, medium confidence
Claude Sonnet 4.657.3%74.6%estimated ± 6.9 pp, medium confidence
Command A+17.6%0.4%estimated ± 6.9 pp, low confidence
Composer 2.558.4%75.9%estimated ± 6.9 pp, medium confidence
Gemini 3.1 Flash-Lite34.1%16.9%estimated ± 6.9 pp, low confidence
Gemini 3.1 Pro70.8%84.2%estimated ± 6.9 pp, medium confidence
Gemini 3.6 Flash73.8%85.2%estimated ± 6.9 pp, medium confidence
Gemini 3 Flash53.9%69.7%estimated ± 6.9 pp, medium confidence
GLM-5.156.9%74.1%estimated ± 6.9 pp, medium confidence
GPT-5.4 mini54.7%71.0%estimated ± 6.9 pp, medium confidence
GPT-5.4 nano41.6%38.8%estimated ± 6.9 pp, low confidence
GPT-5.576.4%85.8%estimated ± 6.9 pp, medium confidence
Grok 4.2044.2%46.8%estimated ± 6.9 pp, low confidence
Grok 4.341.9%39.7%estimated ± 6.9 pp, low confidence
Kimi K2.653.6%69.2%estimated ± 6.9 pp, medium confidence
Kimi K2.7 Code67.0%82.6%estimated ± 6.9 pp, medium confidence
Laguna M.134.1%16.9%estimated ± 6.9 pp, low confidence
Laguna XS.225.8%3.8%estimated ± 6.9 pp, low confidence
Mercury 2.534.5%17.9%estimated ± 6.9 pp, low confidence
MiMo-V2.560.7%78.3%estimated ± 6.9 pp, medium confidence
MiMo-V2.5-Pro57.3%74.6%estimated ± 6.9 pp, medium confidence
MiniMax M2.748.7%59.1%estimated ± 6.9 pp, medium confidence
Mistral Medium 3.5 128B39.0%30.6%estimated ± 6.9 pp, low confidence
Qwen3.6 Plus53.2%68.5%estimated ± 6.9 pp, medium confidence
Qwen3.7 Max61.0%78.5%estimated ± 6.9 pp, medium confidence
Qwen3.7 Plus52.8%67.8%estimated ± 6.9 pp, medium confidence