benchgap
Calibration

Terminal-Bench 2.1 (Vals AI) → AA Terminal-Bench 4.0

AA Terminal-Bench 4.0 is estimated from Terminal-Bench 2.1 (Vals AI) with a linear curve fitted on 13 models measured on both: y = 1.1299·x + -0.5411, R² = 0.59, cross-validated error 14.0 pp. It is used for 29 estimates.

Estimated modelTerminal-Bench 2.1 (Vals AI)AA Terminal-Bench 4.0Source
Claude Haiku 4.543.8%0.0%estimated ± 14.0 pp, low confidence
Claude Opus 4.768.5%23.3%estimated ± 14.0 pp, low confidence
Claude Opus 584.6%41.5%estimated ± 14.0 pp, low confidence
Claude Sonnet 4.657.3%10.6%estimated ± 14.0 pp, low confidence
Command A+17.6%0.0%estimated ± 14.0 pp, low confidence
Composer 2.558.4%11.9%estimated ± 14.0 pp, low confidence
Gemini 3.1 Flash-Lite34.1%0.0%estimated ± 14.0 pp, low confidence
Gemini 3.1 Pro70.8%25.9%estimated ± 14.0 pp, low confidence
Gemini 3.6 Flash73.8%29.3%estimated ± 14.0 pp, low confidence
Gemini 3 Flash53.9%6.8%estimated ± 14.0 pp, low confidence
GLM-5.156.9%10.2%estimated ± 14.0 pp, low confidence
GPT-5.4 mini54.7%7.7%estimated ± 14.0 pp, low confidence
GPT-5.4 nano41.6%0.0%estimated ± 14.0 pp, low confidence
GPT-5.576.4%32.2%estimated ± 14.0 pp, low confidence
Grok 4.2044.2%0.0%estimated ± 14.0 pp, low confidence
Grok 4.341.9%0.0%estimated ± 14.0 pp, low confidence
Grok 4.678.3%34.4%estimated ± 14.0 pp, low confidence
Kimi K2.653.6%6.5%estimated ± 14.0 pp, low confidence
Kimi K2.7 Code67.0%21.6%estimated ± 14.0 pp, low confidence
Laguna M.134.1%0.0%estimated ± 14.0 pp, low confidence
Laguna XS.225.8%0.0%estimated ± 14.0 pp, low confidence
Mercury 2.534.5%0.0%estimated ± 14.0 pp, low confidence
MiMo-V2.560.7%14.5%estimated ± 14.0 pp, low confidence
MiMo-V2.5-Pro57.3%10.6%estimated ± 14.0 pp, low confidence
MiniMax M2.748.7%0.9%estimated ± 14.0 pp, low confidence
Mistral Medium 3.5 128B39.0%0.0%estimated ± 14.0 pp, low confidence
Qwen3.6 Plus53.2%6.0%estimated ± 14.0 pp, low confidence
Qwen3.7 Max61.0%14.8%estimated ± 14.0 pp, low confidence
Qwen3.7 Plus52.8%5.6%estimated ± 14.0 pp, low confidence