benchgap
Calibration

Terminal-Bench 2.1 (Vals AI) → Terminal-Bench 2.0

Terminal-Bench 2.0 is estimated from Terminal-Bench 2.1 (Vals AI) with a Michaelis–Menten curve fitted on 18 models measured on both: y = 1.9912·x / (1.19573 + x), R² = 0.76, cross-validated error 5.9 pp. It is used for 43 estimates.

Estimated modelTerminal-Bench 2.1 (Vals AI)Terminal-Bench 2.0Source
Claude Fable 580.5%80.1%estimated ± 5.9 pp, low confidence
Claude Fable 5.185.0%82.7%estimated ± 5.9 pp, low confidence
Claude Haiku 4.543.8%53.4%estimated ± 5.9 pp, medium confidence
Claude Opus 4.768.5%72.5%estimated ± 5.9 pp, medium confidence
Claude Opus 4.871.9%74.8%estimated ± 5.9 pp, medium confidence
Claude Opus 584.6%82.5%estimated ± 5.9 pp, low confidence
Claude Sonnet 574.5%76.4%estimated ± 5.9 pp, medium confidence
Command A+17.6%25.5%estimated ± 5.9 pp, low confidence
DeepSeek V4.1 Flash74.5%76.4%estimated ± 5.9 pp, medium confidence
Gemini 3.1 Flash-Lite34.1%44.2%estimated ± 5.9 pp, medium confidence
Gemini 3.1 Pro70.8%74.1%estimated ± 5.9 pp, medium confidence
Gemini 3.5 Flash74.2%76.2%estimated ± 5.9 pp, medium confidence
Gemini 3.5 Flash-Lite50.2%58.9%estimated ± 5.9 pp, medium confidence
Gemini 3.6 Flash73.8%76.0%estimated ± 5.9 pp, medium confidence
Gemini 3.7 Flash77.5%78.3%estimated ± 5.9 pp, low confidence
Gemini 3.8 Flash81.3%80.6%estimated ± 5.9 pp, low confidence
Gemini 3 Flash53.9%61.9%estimated ± 5.9 pp, medium confidence
GLM-5.267.8%72.1%estimated ± 5.9 pp, medium confidence
GLM-5.371.5%74.5%estimated ± 5.9 pp, medium confidence
GLM-5.3-Flash62.9%68.6%estimated ± 5.9 pp, medium confidence
GPT-5.6 Luna79.0%79.2%estimated ± 5.9 pp, low confidence
GPT-5.6 Sol85.8%83.2%estimated ± 5.9 pp, low confidence
GPT-5.6 Terra77.5%78.3%estimated ± 5.9 pp, low confidence
GPT-6 Astra87.3%84.0%estimated ± 5.9 pp, low confidence
Grok 4.341.9%51.7%estimated ± 5.9 pp, medium confidence
Grok 4.567.8%72.1%estimated ± 5.9 pp, medium confidence
Grok 4.678.3%78.8%estimated ± 5.9 pp, low confidence
Grok 4.773.4%75.7%estimated ± 5.9 pp, medium confidence
Hy4 preview55.1%62.8%estimated ± 5.9 pp, medium confidence
Inkling47.6%56.7%estimated ± 5.9 pp, medium confidence
Inkling-Small55.1%62.8%estimated ± 5.9 pp, medium confidence
Kimi K2.7 Code67.0%71.5%estimated ± 5.9 pp, medium confidence
Kimi K380.9%80.4%estimated ± 5.9 pp, low confidence
Ling 3.0 Flash50.2%58.9%estimated ± 5.9 pp, medium confidence
Mercury 2.534.5%44.6%estimated ± 5.9 pp, medium confidence
MiniMax M353.6%61.6%estimated ± 5.9 pp, medium confidence
Mistral Medium 3.5 128B39.0%49.0%estimated ± 5.9 pp, medium confidence
Muse Spark 1.169.3%73.1%estimated ± 5.9 pp, medium confidence
Muse Spark 1.269.7%73.3%estimated ± 5.9 pp, medium confidence
Muse Spark 1.372.3%75.0%estimated ± 5.9 pp, medium confidence
Nemotron 3 Ultra50.9%59.5%estimated ± 5.9 pp, medium confidence
Qwen3.8-27B58.4%65.3%estimated ± 5.9 pp, medium confidence
Qwen3.8 Max67.4%71.8%estimated ± 5.9 pp, medium confidence