Calibration
Terminal-Bench 2.1 (Vals AI) → Terminal-Bench 3.0
Terminal-Bench 3.0 is estimated from Terminal-Bench 2.1 (Vals AI) with a Hill curve fitted on 11 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.98911^6.00 + x^6.00), R² = 0.67, cross-validated error 6.9 pp. It is used for 50 estimates.
| Estimated model | Terminal-Bench 2.1 (Vals AI) | Terminal-Bench 3.0 | Source |
|---|---|---|---|
| Claude Fable 5.1 | 85.0% | 34.5% | estimated ± 6.9 pp, medium confidence |
| Claude Haiku 4.5 | 43.8% | 0.9% | estimated ± 6.9 pp, low confidence |
| Claude Opus 4.7 | 68.5% | 11.9% | estimated ± 6.9 pp, medium confidence |
| Claude Sonnet 4.6 | 57.3% | 4.4% | estimated ± 6.9 pp, low confidence |
| Command A+ | 17.6% | 0.0% | estimated ± 6.9 pp, low confidence |
| Composer 2.5 | 58.4% | 4.9% | estimated ± 6.9 pp, low confidence |
| DeepSeek V4.1 Flash | 74.5% | 18.5% | estimated ± 6.9 pp, medium confidence |
| DeepSeek V4 Flash 0731 | 67.0% | 10.6% | estimated ± 6.9 pp, low confidence |
| DeepSeek V4 Pro 0813 | 54.7% | 3.3% | estimated ± 6.9 pp, low confidence |
| Gemini 3.1 Flash-Lite | 34.1% | 0.2% | estimated ± 6.9 pp, low confidence |
| Gemini 3.1 Pro | 70.8% | 14.2% | estimated ± 6.9 pp, medium confidence |
| Gemini 3.5 Flash | 74.2% | 18.2% | estimated ± 6.9 pp, medium confidence |
| Gemini 3.5 Flash-Lite | 50.2% | 2.0% | estimated ± 6.9 pp, low confidence |
| Gemini 3.6 Flash | 73.8% | 17.7% | estimated ± 6.9 pp, medium confidence |
| Gemini 3.8 Flash | 81.3% | 28.3% | estimated ± 6.9 pp, medium confidence |
| Gemini 3 Flash | 53.9% | 3.1% | estimated ± 6.9 pp, low confidence |
| GLM-5.1 | 56.9% | 4.2% | estimated ± 6.9 pp, low confidence |
| GLM-5.3 | 71.5% | 15.0% | estimated ± 6.9 pp, medium confidence |
| GLM-5.3-Flash | 62.9% | 7.4% | estimated ± 6.9 pp, low confidence |
| GPT-5.4 mini | 54.7% | 3.3% | estimated ± 6.9 pp, low confidence |
| GPT-5.4 nano | 41.6% | 0.7% | estimated ± 6.9 pp, low confidence |
| GPT-5.5 | 76.4% | 21.0% | estimated ± 6.9 pp, medium confidence |
| GPT-6 Astra | 87.3% | 38.5% | estimated ± 6.9 pp, low confidence |
| Grok 4.20 | 44.2% | 0.9% | estimated ± 6.9 pp, low confidence |
| Grok 4.3 | 41.9% | 0.7% | estimated ± 6.9 pp, low confidence |
| Grok 4.7 | 73.4% | 17.2% | estimated ± 6.9 pp, medium confidence |
| Hy4 preview | 55.1% | 3.5% | estimated ± 6.9 pp, low confidence |
| Inkling | 47.6% | 1.5% | estimated ± 6.9 pp, low confidence |
| Inkling-Small | 55.1% | 3.5% | estimated ± 6.9 pp, low confidence |
| Kimi K2.6 | 53.6% | 3.0% | estimated ± 6.9 pp, low confidence |
| Kimi K2.7 Code | 67.0% | 10.6% | estimated ± 6.9 pp, low confidence |
| Kimi K3 | 80.9% | 27.6% | estimated ± 6.9 pp, medium confidence |
| Laguna M.1 | 34.1% | 0.2% | estimated ± 6.9 pp, low confidence |
| Laguna XS.2 | 25.8% | 0.0% | estimated ± 6.9 pp, low confidence |
| Ling 3.0 Flash | 50.2% | 2.0% | estimated ± 6.9 pp, low confidence |
| Mercury 2.5 | 34.5% | 0.2% | estimated ± 6.9 pp, low confidence |
| MiMo-V2.5 | 60.7% | 6.1% | estimated ± 6.9 pp, low confidence |
| MiMo-V2.5-Pro | 57.3% | 4.4% | estimated ± 6.9 pp, low confidence |
| MiniMax M2.7 | 48.7% | 1.7% | estimated ± 6.9 pp, low confidence |
| MiniMax M3 | 53.6% | 3.0% | estimated ± 6.9 pp, low confidence |
| Mistral Medium 3.5 128B | 39.0% | 0.4% | estimated ± 6.9 pp, low confidence |
| Muse Spark 1.1 | 69.3% | 12.7% | estimated ± 6.9 pp, medium confidence |
| Muse Spark 1.2 | 69.7% | 13.1% | estimated ± 6.9 pp, medium confidence |
| Muse Spark 1.3 | 72.3% | 15.9% | estimated ± 6.9 pp, medium confidence |
| Nemotron 3 Ultra | 50.9% | 2.2% | estimated ± 6.9 pp, low confidence |
| Qwen3.6 Plus | 53.2% | 2.8% | estimated ± 6.9 pp, low confidence |
| Qwen3.7 Max | 61.0% | 6.3% | estimated ± 6.9 pp, low confidence |
| Qwen3.7 Plus | 52.8% | 2.7% | estimated ± 6.9 pp, low confidence |
| Qwen3.8-27B | 58.4% | 4.9% | estimated ± 6.9 pp, low confidence |
| Qwen3.8 Max | 67.4% | 10.9% | estimated ± 6.9 pp, low confidence |