benchgap
Calibration

AA Terminal-Bench 4.0 → AA Terminal-Bench 2.1

AA Terminal-Bench 2.1 is estimated from AA Terminal-Bench 4.0 with a Hill curve fitted on 12 models measured on both: y = 0.5090 + (0.8694 − 0.5090)·x^1.86 / (0.02598^1.86 + x^1.86), R² = 0.98, cross-validated error 2.6 pp. It is used for 12 estimates.

Estimated modelAA Terminal-Bench 4.0AA Terminal-Bench 2.1Source
Claude Haiku 5.532.8%86.6%estimated ± 2.6 pp, high confidence
Claude Opus 5.559.6%86.8%estimated ± 2.6 pp, medium confidence
Claude Sonnet 5.563.6%86.8%estimated ± 2.6 pp, medium confidence
DeepSeek V4.1 Flash26.8%86.5%estimated ± 2.6 pp, high confidence
Gemini 4 Argon57.1%86.8%estimated ± 2.6 pp, high confidence
GPT-6.1 Sol56.1%86.8%estimated ± 2.6 pp, high confidence
GPT-6 Luna12.6%85.1%estimated ± 2.6 pp, high confidence
GPT-6 Sol43.9%86.8%estimated ± 2.6 pp, high confidence
Grok 4.725.8%86.4%estimated ± 2.6 pp, high confidence
MiMo-V2.6-Pro34.8%86.7%estimated ± 2.6 pp, high confidence
Mistral Large 426.8%86.5%estimated ± 2.6 pp, high confidence
Step 5 Preview33.3%86.6%estimated ± 2.6 pp, high confidence