Calibration
AA Terminal-Bench 4.0 → Terminal-Bench-Science 0.1
Terminal-Bench-Science 0.1 is estimated from AA Terminal-Bench 4.0 with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.6668 − 0.0000)·x^6.00 / (0.41261^6.00 + x^6.00), R² = 0.60, cross-validated error 4.1 pp. It is used for 19 estimates.
| Estimated model | AA Terminal-Bench 4.0 | Terminal-Bench-Science 0.1 | Source |
|---|---|---|---|
| Claude Haiku 5.5 | 32.8% | 13.4% | estimated ± 4.1 pp, low confidence |
| DeepSeek V4.1 Flash | 26.8% | 4.7% | estimated ± 4.1 pp, low confidence |
| Gemini 3.8 Flash | 19.7% | 0.8% | estimated ± 4.1 pp, low confidence |
| Gemini 4 Argon | 57.1% | 58.4% | estimated ± 4.1 pp, medium confidence |
| GLM-5.3 | 41.9% | 34.9% | estimated ± 4.1 pp, low confidence |
| GLM-5.3-Flash | 32.8% | 13.4% | estimated ± 4.1 pp, low confidence |
| GPT-6 Luna | 12.6% | 0.1% | estimated ± 4.1 pp, low confidence |
| GPT-6 Sol | 43.9% | 39.5% | estimated ± 4.1 pp, low confidence |
| Grok 4.7 | 25.8% | 3.8% | estimated ± 4.1 pp, low confidence |
| Inkling | 1.0% | 0.0% | estimated ± 4.1 pp, low confidence |
| Kimi K3 | 12.6% | 0.1% | estimated ± 4.1 pp, low confidence |
| MiMo-V2.6-Pro | 34.8% | 17.6% | estimated ± 4.1 pp, low confidence |
| MiniMax M3 | 2.0% | 0.0% | estimated ± 4.1 pp, low confidence |
| Mistral Large 4 | 26.8% | 4.7% | estimated ± 4.1 pp, low confidence |
| Muse Glimmer 30B | 0.5% | 0.0% | estimated ± 4.1 pp, low confidence |
| Muse Spark 1.3 | 33.3% | 14.4% | estimated ± 4.1 pp, low confidence |
| Nemotron 3 Ultra | 0.5% | 0.0% | estimated ± 4.1 pp, low confidence |
| Qwen3.8-27B | 5.6% | 0.0% | estimated ± 4.1 pp, low confidence |
| Step 5 Preview | 33.3% | 14.4% | estimated ± 4.1 pp, low confidence |