Calibration
Terminal-Bench 4.0 → Terminal-Bench 2.1 (Vals AI)
Terminal-Bench 2.1 (Vals AI) is estimated from Terminal-Bench 4.0 with a Hill curve fitted on 5 models measured on both: y = 0.7610 + (1.2000 − 0.7610)·x^6.00 / (0.70056^6.00 + x^6.00), R² = 0.70, cross-validated error 4.6 pp. It is used for 12 estimates.
| Estimated model | Terminal-Bench 4.0 | Terminal-Bench 2.1 (Vals AI) | Source |
|---|---|---|---|
| Claude Haiku 5.5 | 39.2% | 77.4% | estimated ± 4.6 pp, medium confidence |
| Claude Mythos 5.1 | 60.9% | 89.3% | estimated ± 4.6 pp, low confidence |
| Claude Opus 5.5 | 66.4% | 94.5% | estimated ± 4.6 pp, low confidence |
| Claude Sonnet 5.5 | 70.6% | 98.6% | estimated ± 4.6 pp, low confidence |
| Gemini 4 Argon | 57.4% | 86.3% | estimated ± 4.6 pp, medium confidence |
| Ling 3.1 Flash | 40.4% | 77.7% | estimated ± 4.6 pp, medium confidence |
| MiMo-V2.6-Flash | 28.8% | 76.3% | estimated ± 4.6 pp, medium confidence |
| MiMo-V2.6-Pro | 34.9% | 76.8% | estimated ± 4.6 pp, medium confidence |
| Pareto 26.10 Preview | 50.8% | 81.7% | estimated ± 4.6 pp, medium confidence |
| Pareto 26.9 | 51.0% | 81.8% | estimated ± 4.6 pp, medium confidence |
| Step 5 Preview | 33.3% | 76.6% | estimated ± 4.6 pp, medium confidence |
| SWE-2 | 27.3% | 76.2% | estimated ± 4.6 pp, medium confidence |