Calibration
AA Terminal-Bench 4.0 → Terminal-Bench 2.1
Terminal-Bench 2.1 is estimated from AA Terminal-Bench 4.0 with a Michaelis–Menten + offset curve fitted on 13 models measured on both: y = 0.4818 + 0.4344·x / (0.02665 + x), R² = 0.95, cross-validated error 3.9 pp. It is used for 11 estimates.
| Estimated model | AA Terminal-Bench 4.0 | Terminal-Bench 2.1 | Source |
|---|---|---|---|
| Claude Fable 5.1 | 52.0% | 89.5% | estimated ± 3.9 pp, medium confidence |
| Claude Haiku 5.5 | 32.8% | 88.4% | estimated ± 3.9 pp, high confidence |
| Claude Opus 5.5 | 59.6% | 89.8% | estimated ± 3.9 pp, medium confidence |
| Claude Sonnet 5.5 | 63.6% | 89.9% | estimated ± 3.9 pp, medium confidence |
| Gemini 4 Argon | 57.1% | 89.7% | estimated ± 3.9 pp, medium confidence |
| GPT-6.1 Sol | 56.1% | 89.6% | estimated ± 3.9 pp, medium confidence |
| GPT-6 Astra | 59.1% | 89.7% | estimated ± 3.9 pp, medium confidence |
| GPT-6 Luna | 12.6% | 84.0% | estimated ± 3.9 pp, high confidence |
| GPT-6 Sol | 43.9% | 89.1% | estimated ± 3.9 pp, medium confidence |
| Grok 4.7 | 25.8% | 87.6% | estimated ± 3.9 pp, high confidence |
| Mistral Large 4 | 26.8% | 87.7% | estimated ± 3.9 pp, high confidence |