benchgap
Calibration

AA Terminal-Bench 4.0 → Terminal-Bench-Science 0.1

Terminal-Bench-Science 0.1 is estimated from AA Terminal-Bench 4.0 with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.6668 − 0.0000)·x^6.00 / (0.41261^6.00 + x^6.00), R² = 0.60, cross-validated error 4.1 pp. It is used for 19 estimates.

Estimated modelAA Terminal-Bench 4.0Terminal-Bench-Science 0.1Source
Claude Haiku 5.532.8%13.4%estimated ± 4.1 pp, low confidence
DeepSeek V4.1 Flash26.8%4.7%estimated ± 4.1 pp, low confidence
Gemini 3.8 Flash19.7%0.8%estimated ± 4.1 pp, low confidence
Gemini 4 Argon57.1%58.4%estimated ± 4.1 pp, medium confidence
GLM-5.341.9%34.9%estimated ± 4.1 pp, low confidence
GLM-5.3-Flash32.8%13.4%estimated ± 4.1 pp, low confidence
GPT-6 Luna12.6%0.1%estimated ± 4.1 pp, low confidence
GPT-6 Sol43.9%39.5%estimated ± 4.1 pp, low confidence
Grok 4.725.8%3.8%estimated ± 4.1 pp, low confidence
Inkling1.0%0.0%estimated ± 4.1 pp, low confidence
Kimi K312.6%0.1%estimated ± 4.1 pp, low confidence
MiMo-V2.6-Pro34.8%17.6%estimated ± 4.1 pp, low confidence
MiniMax M32.0%0.0%estimated ± 4.1 pp, low confidence
Mistral Large 426.8%4.7%estimated ± 4.1 pp, low confidence
Muse Glimmer 30B0.5%0.0%estimated ± 4.1 pp, low confidence
Muse Spark 1.333.3%14.4%estimated ± 4.1 pp, low confidence
Nemotron 3 Ultra0.5%0.0%estimated ± 4.1 pp, low confidence
Qwen3.8-27B5.6%0.0%estimated ± 4.1 pp, low confidence
Step 5 Preview33.3%14.4%estimated ± 4.1 pp, low confidence