benchgap
Calibration

Terminal-Bench 2.1 (Vals AI) → AA Terminal-Bench 2.1

AA Terminal-Bench 2.1 is estimated from Terminal-Bench 2.1 (Vals AI) with a Hill curve fitted on 11 models measured on both: y = 0.0000 + (0.9086 − 0.0000)·x^6.00 / (0.45312^6.00 + x^6.00), R² = 0.93, cross-validated error 4.4 pp. It is used for 29 estimates.

Estimated modelTerminal-Bench 2.1 (Vals AI)AA Terminal-Bench 2.1Source
Claude Haiku 4.543.8%40.8%estimated ± 4.4 pp, medium confidence
Claude Opus 4.768.5%83.8%estimated ± 4.4 pp, high confidence
Claude Opus 584.6%88.8%estimated ± 4.4 pp, high confidence
Claude Sonnet 4.657.3%73.0%estimated ± 4.4 pp, high confidence
Command A+17.6%0.3%estimated ± 4.4 pp, medium confidence
Composer 2.558.4%74.6%estimated ± 4.4 pp, high confidence
Gemini 3.1 Flash-Lite34.1%14.0%estimated ± 4.4 pp, medium confidence
Gemini 3.1 Pro70.8%85.0%estimated ± 4.4 pp, high confidence
Gemini 3.6 Flash73.8%86.2%estimated ± 4.4 pp, high confidence
Gemini 3 Flash53.9%67.2%estimated ± 4.4 pp, high confidence
GLM-5.156.9%72.4%estimated ± 4.4 pp, high confidence
GPT-5.4 mini54.7%68.7%estimated ± 4.4 pp, high confidence
GPT-5.4 nano41.6%34.0%estimated ± 4.4 pp, medium confidence
GPT-5.576.4%87.1%estimated ± 4.4 pp, high confidence
Grok 4.2044.2%42.1%estimated ± 4.4 pp, medium confidence
Grok 4.341.9%35.0%estimated ± 4.4 pp, medium confidence
Grok 4.678.3%87.6%estimated ± 4.4 pp, high confidence
Kimi K2.653.6%66.6%estimated ± 4.4 pp, high confidence
Kimi K2.7 Code67.0%82.9%estimated ± 4.4 pp, high confidence
Laguna M.134.1%14.0%estimated ± 4.4 pp, medium confidence
Laguna XS.225.8%3.0%estimated ± 4.4 pp, medium confidence
Mercury 2.534.5%14.8%estimated ± 4.4 pp, medium confidence
MiMo-V2.560.7%77.5%estimated ± 4.4 pp, high confidence
MiMo-V2.5-Pro57.3%73.0%estimated ± 4.4 pp, high confidence
MiniMax M2.748.7%55.1%estimated ± 4.4 pp, high confidence
Mistral Medium 3.5 128B39.0%26.3%estimated ± 4.4 pp, medium confidence
Qwen3.6 Plus53.2%65.8%estimated ± 4.4 pp, high confidence
Qwen3.7 Max61.0%77.8%estimated ± 4.4 pp, high confidence
Qwen3.7 Plus52.8%64.9%estimated ± 4.4 pp, high confidence