Calibration
Terminal-Bench 2.1 (Vals AI) → AA Terminal-Bench 4.0
AA Terminal-Bench 4.0 is estimated from Terminal-Bench 2.1 (Vals AI) with a linear curve fitted on 13 models measured on both: y = 1.1299·x + -0.5411, R² = 0.59, cross-validated error 14.0 pp. It is used for 29 estimates.
| Estimated model | Terminal-Bench 2.1 (Vals AI) | AA Terminal-Bench 4.0 | Source |
|---|---|---|---|
| Claude Haiku 4.5 | 43.8% | 0.0% | estimated ± 14.0 pp, low confidence |
| Claude Opus 4.7 | 68.5% | 23.3% | estimated ± 14.0 pp, low confidence |
| Claude Opus 5 | 84.6% | 41.5% | estimated ± 14.0 pp, low confidence |
| Claude Sonnet 4.6 | 57.3% | 10.6% | estimated ± 14.0 pp, low confidence |
| Command A+ | 17.6% | 0.0% | estimated ± 14.0 pp, low confidence |
| Composer 2.5 | 58.4% | 11.9% | estimated ± 14.0 pp, low confidence |
| Gemini 3.1 Flash-Lite | 34.1% | 0.0% | estimated ± 14.0 pp, low confidence |
| Gemini 3.1 Pro | 70.8% | 25.9% | estimated ± 14.0 pp, low confidence |
| Gemini 3.6 Flash | 73.8% | 29.3% | estimated ± 14.0 pp, low confidence |
| Gemini 3 Flash | 53.9% | 6.8% | estimated ± 14.0 pp, low confidence |
| GLM-5.1 | 56.9% | 10.2% | estimated ± 14.0 pp, low confidence |
| GPT-5.4 mini | 54.7% | 7.7% | estimated ± 14.0 pp, low confidence |
| GPT-5.4 nano | 41.6% | 0.0% | estimated ± 14.0 pp, low confidence |
| GPT-5.5 | 76.4% | 32.2% | estimated ± 14.0 pp, low confidence |
| Grok 4.20 | 44.2% | 0.0% | estimated ± 14.0 pp, low confidence |
| Grok 4.3 | 41.9% | 0.0% | estimated ± 14.0 pp, low confidence |
| Grok 4.6 | 78.3% | 34.4% | estimated ± 14.0 pp, low confidence |
| Kimi K2.6 | 53.6% | 6.5% | estimated ± 14.0 pp, low confidence |
| Kimi K2.7 Code | 67.0% | 21.6% | estimated ± 14.0 pp, low confidence |
| Laguna M.1 | 34.1% | 0.0% | estimated ± 14.0 pp, low confidence |
| Laguna XS.2 | 25.8% | 0.0% | estimated ± 14.0 pp, low confidence |
| Mercury 2.5 | 34.5% | 0.0% | estimated ± 14.0 pp, low confidence |
| MiMo-V2.5 | 60.7% | 14.5% | estimated ± 14.0 pp, low confidence |
| MiMo-V2.5-Pro | 57.3% | 10.6% | estimated ± 14.0 pp, low confidence |
| MiniMax M2.7 | 48.7% | 0.9% | estimated ± 14.0 pp, low confidence |
| Mistral Medium 3.5 128B | 39.0% | 0.0% | estimated ± 14.0 pp, low confidence |
| Qwen3.6 Plus | 53.2% | 6.0% | estimated ± 14.0 pp, low confidence |
| Qwen3.7 Max | 61.0% | 14.8% | estimated ± 14.0 pp, low confidence |
| Qwen3.7 Plus | 52.8% | 5.6% | estimated ± 14.0 pp, low confidence |