benchgap
Calibration

Terminal-Bench 2.1 (Vals AI) → Terminal-Bench 3.0

Terminal-Bench 3.0 is estimated from Terminal-Bench 2.1 (Vals AI) with a Hill curve fitted on 11 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.98911^6.00 + x^6.00), R² = 0.67, cross-validated error 6.9 pp. It is used for 50 estimates.

Estimated modelTerminal-Bench 2.1 (Vals AI)Terminal-Bench 3.0Source
Claude Fable 5.185.0%34.5%estimated ± 6.9 pp, medium confidence
Claude Haiku 4.543.8%0.9%estimated ± 6.9 pp, low confidence
Claude Opus 4.768.5%11.9%estimated ± 6.9 pp, medium confidence
Claude Sonnet 4.657.3%4.4%estimated ± 6.9 pp, low confidence
Command A+17.6%0.0%estimated ± 6.9 pp, low confidence
Composer 2.558.4%4.9%estimated ± 6.9 pp, low confidence
DeepSeek V4.1 Flash74.5%18.5%estimated ± 6.9 pp, medium confidence
DeepSeek V4 Flash 073167.0%10.6%estimated ± 6.9 pp, low confidence
DeepSeek V4 Pro 081354.7%3.3%estimated ± 6.9 pp, low confidence
Gemini 3.1 Flash-Lite34.1%0.2%estimated ± 6.9 pp, low confidence
Gemini 3.1 Pro70.8%14.2%estimated ± 6.9 pp, medium confidence
Gemini 3.5 Flash74.2%18.2%estimated ± 6.9 pp, medium confidence
Gemini 3.5 Flash-Lite50.2%2.0%estimated ± 6.9 pp, low confidence
Gemini 3.6 Flash73.8%17.7%estimated ± 6.9 pp, medium confidence
Gemini 3.8 Flash81.3%28.3%estimated ± 6.9 pp, medium confidence
Gemini 3 Flash53.9%3.1%estimated ± 6.9 pp, low confidence
GLM-5.156.9%4.2%estimated ± 6.9 pp, low confidence
GLM-5.371.5%15.0%estimated ± 6.9 pp, medium confidence
GLM-5.3-Flash62.9%7.4%estimated ± 6.9 pp, low confidence
GPT-5.4 mini54.7%3.3%estimated ± 6.9 pp, low confidence
GPT-5.4 nano41.6%0.7%estimated ± 6.9 pp, low confidence
GPT-5.576.4%21.0%estimated ± 6.9 pp, medium confidence
GPT-6 Astra87.3%38.5%estimated ± 6.9 pp, low confidence
Grok 4.2044.2%0.9%estimated ± 6.9 pp, low confidence
Grok 4.341.9%0.7%estimated ± 6.9 pp, low confidence
Grok 4.773.4%17.2%estimated ± 6.9 pp, medium confidence
Hy4 preview55.1%3.5%estimated ± 6.9 pp, low confidence
Inkling47.6%1.5%estimated ± 6.9 pp, low confidence
Inkling-Small55.1%3.5%estimated ± 6.9 pp, low confidence
Kimi K2.653.6%3.0%estimated ± 6.9 pp, low confidence
Kimi K2.7 Code67.0%10.6%estimated ± 6.9 pp, low confidence
Kimi K380.9%27.6%estimated ± 6.9 pp, medium confidence
Laguna M.134.1%0.2%estimated ± 6.9 pp, low confidence
Laguna XS.225.8%0.0%estimated ± 6.9 pp, low confidence
Ling 3.0 Flash50.2%2.0%estimated ± 6.9 pp, low confidence
Mercury 2.534.5%0.2%estimated ± 6.9 pp, low confidence
MiMo-V2.560.7%6.1%estimated ± 6.9 pp, low confidence
MiMo-V2.5-Pro57.3%4.4%estimated ± 6.9 pp, low confidence
MiniMax M2.748.7%1.7%estimated ± 6.9 pp, low confidence
MiniMax M353.6%3.0%estimated ± 6.9 pp, low confidence
Mistral Medium 3.5 128B39.0%0.4%estimated ± 6.9 pp, low confidence
Muse Spark 1.169.3%12.7%estimated ± 6.9 pp, medium confidence
Muse Spark 1.269.7%13.1%estimated ± 6.9 pp, medium confidence
Muse Spark 1.372.3%15.9%estimated ± 6.9 pp, medium confidence
Nemotron 3 Ultra50.9%2.2%estimated ± 6.9 pp, low confidence
Qwen3.6 Plus53.2%2.8%estimated ± 6.9 pp, low confidence
Qwen3.7 Max61.0%6.3%estimated ± 6.9 pp, low confidence
Qwen3.7 Plus52.8%2.7%estimated ± 6.9 pp, low confidence
Qwen3.8-27B58.4%4.9%estimated ± 6.9 pp, low confidence
Qwen3.8 Max67.4%10.9%estimated ± 6.9 pp, low confidence