benchgap
Calibration

Terminal-Bench 4.0 → Terminal-Bench 2.1

Terminal-Bench 2.1 is estimated from Terminal-Bench 4.0 with a Michaelis–Menten + offset curve fitted on 15 models measured on both: y = 0.3902 + 0.5038·x / (0.01405 + x), R² = 0.95, cross-validated error 4.0 pp. It is used for 15 estimates.

Estimated modelTerminal-Bench 4.0Terminal-Bench 2.1Source
Claude Sonnet 5.5 (max with fallback)63.6%88.3%estimated ± 4.0 pp, medium confidence
Claude Opus 5.5 (max with fallback)59.6%88.2%estimated ± 4.0 pp, medium confidence
Claude Opus 5.5 (xhigh with fallback)59.6%88.2%estimated ± 4.0 pp, medium confidence
GPT-6 Astra (xhigh)59.6%88.2%estimated ± 4.0 pp, medium confidence
Gemini 4 Argon (high)57.1%88.2%estimated ± 4.0 pp, high confidence
Claude Sonnet 5.5 (xhigh with fallback)57.1%88.2%estimated ± 4.0 pp, high confidence
Claude Opus 5.5 (high with fallback)56.6%88.2%estimated ± 4.0 pp, high confidence
GPT-6.1 Sol (max)56.1%88.2%estimated ± 4.0 pp, high confidence
GPT-6 Sol (max)43.9%87.8%estimated ± 4.0 pp, high confidence
MiMo-V2.6-Pro34.8%87.4%estimated ± 4.0 pp, high confidence
Step 5 Preview33.3%87.4%estimated ± 4.0 pp, high confidence
DeepSeek V4.1 Flash (max)26.8%86.9%estimated ± 4.0 pp, high confidence
Mistral Large 4 Preview26.8%86.9%estimated ± 4.0 pp, high confidence
Grok 4.7 (xhigh)25.8%86.8%estimated ± 4.0 pp, high confidence
GPT-6 Luna (max)12.6%84.3%estimated ± 4.0 pp, high confidence