Calibration
Terminal-Bench 2.1 (Vals AI) → Terminal-Bench 2.1
Terminal-Bench 2.1 is estimated from Terminal-Bench 2.1 (Vals AI) with a Hill curve fitted on 29 models measured on both: y = 0.0000 + (0.8870 − 0.0000)·x^6.00 / (0.43383^6.00 + x^6.00), R² = 0.67, cross-validated error 6.9 pp. It is used for 27 estimates.
| Estimated model | Terminal-Bench 2.1 (Vals AI) | Terminal-Bench 2.1 | Source |
|---|---|---|---|
| Claude Haiku 4.5 | 43.8% | 45.6% | estimated ± 6.9 pp, low confidence |
| Claude Opus 4.7 | 68.5% | 83.3% | estimated ± 6.9 pp, medium confidence |
| Claude Sonnet 4.6 | 57.3% | 74.6% | estimated ± 6.9 pp, medium confidence |
| Command A+ | 17.6% | 0.4% | estimated ± 6.9 pp, low confidence |
| Composer 2.5 | 58.4% | 75.9% | estimated ± 6.9 pp, medium confidence |
| Gemini 3.1 Flash-Lite | 34.1% | 16.9% | estimated ± 6.9 pp, low confidence |
| Gemini 3.1 Pro | 70.8% | 84.2% | estimated ± 6.9 pp, medium confidence |
| Gemini 3.6 Flash | 73.8% | 85.2% | estimated ± 6.9 pp, medium confidence |
| Gemini 3 Flash | 53.9% | 69.7% | estimated ± 6.9 pp, medium confidence |
| GLM-5.1 | 56.9% | 74.1% | estimated ± 6.9 pp, medium confidence |
| GPT-5.4 mini | 54.7% | 71.0% | estimated ± 6.9 pp, medium confidence |
| GPT-5.4 nano | 41.6% | 38.8% | estimated ± 6.9 pp, low confidence |
| GPT-5.5 | 76.4% | 85.8% | estimated ± 6.9 pp, medium confidence |
| Grok 4.20 | 44.2% | 46.8% | estimated ± 6.9 pp, low confidence |
| Grok 4.3 | 41.9% | 39.7% | estimated ± 6.9 pp, low confidence |
| Kimi K2.6 | 53.6% | 69.2% | estimated ± 6.9 pp, medium confidence |
| Kimi K2.7 Code | 67.0% | 82.6% | estimated ± 6.9 pp, medium confidence |
| Laguna M.1 | 34.1% | 16.9% | estimated ± 6.9 pp, low confidence |
| Laguna XS.2 | 25.8% | 3.8% | estimated ± 6.9 pp, low confidence |
| Mercury 2.5 | 34.5% | 17.9% | estimated ± 6.9 pp, low confidence |
| MiMo-V2.5 | 60.7% | 78.3% | estimated ± 6.9 pp, medium confidence |
| MiMo-V2.5-Pro | 57.3% | 74.6% | estimated ± 6.9 pp, medium confidence |
| MiniMax M2.7 | 48.7% | 59.1% | estimated ± 6.9 pp, medium confidence |
| Mistral Medium 3.5 128B | 39.0% | 30.6% | estimated ± 6.9 pp, low confidence |
| Qwen3.6 Plus | 53.2% | 68.5% | estimated ± 6.9 pp, medium confidence |
| Qwen3.7 Max | 61.0% | 78.5% | estimated ± 6.9 pp, medium confidence |
| Qwen3.7 Plus | 52.8% | 67.8% | estimated ± 6.9 pp, medium confidence |