Calibration
AA Terminal-Bench 4.0 → AA Terminal-Bench 2.1
AA Terminal-Bench 2.1 is estimated from AA Terminal-Bench 4.0 with a Hill curve fitted on 12 models measured on both: y = 0.5090 + (0.8694 − 0.5090)·x^1.86 / (0.02598^1.86 + x^1.86), R² = 0.98, cross-validated error 2.6 pp. It is used for 12 estimates.
| Estimated model | AA Terminal-Bench 4.0 | AA Terminal-Bench 2.1 | Source |
|---|---|---|---|
| Claude Haiku 5.5 | 32.8% | 86.6% | estimated ± 2.6 pp, high confidence |
| Claude Opus 5.5 | 59.6% | 86.8% | estimated ± 2.6 pp, medium confidence |
| Claude Sonnet 5.5 | 63.6% | 86.8% | estimated ± 2.6 pp, medium confidence |
| DeepSeek V4.1 Flash | 26.8% | 86.5% | estimated ± 2.6 pp, high confidence |
| Gemini 4 Argon | 57.1% | 86.8% | estimated ± 2.6 pp, high confidence |
| GPT-6.1 Sol | 56.1% | 86.8% | estimated ± 2.6 pp, high confidence |
| GPT-6 Luna | 12.6% | 85.1% | estimated ± 2.6 pp, high confidence |
| GPT-6 Sol | 43.9% | 86.8% | estimated ± 2.6 pp, high confidence |
| Grok 4.7 | 25.8% | 86.4% | estimated ± 2.6 pp, high confidence |
| MiMo-V2.6-Pro | 34.8% | 86.7% | estimated ± 2.6 pp, high confidence |
| Mistral Large 4 | 26.8% | 86.5% | estimated ± 2.6 pp, high confidence |
| Step 5 Preview | 33.3% | 86.6% | estimated ± 2.6 pp, high confidence |