benchgap
Calibration

AA Terminal-Bench 4.0 → Terminal-Bench 4.0

Terminal-Bench 4.0 is estimated from AA Terminal-Bench 4.0 with a linear curve fitted on 11 models measured on both: y = 0.9896·x + 0.0400, R² = 0.93, cross-validated error 5.1 pp. It is used for 13 estimates.

Estimated modelAA Terminal-Bench 4.0Terminal-Bench 4.0Source
GLM-5.341.9%45.5%estimated ± 5.1 pp, medium confidence
GLM-5.3-Flash32.8%36.5%estimated ± 5.1 pp, medium confidence
GPT-6.1 Sol56.1%59.5%estimated ± 5.1 pp, medium confidence
GPT-6 Luna12.6%16.5%estimated ± 5.1 pp, low confidence
GPT-6 Sol43.9%47.4%estimated ± 5.1 pp, medium confidence
Inkling1.0%5.0%estimated ± 5.1 pp, low confidence
Kimi K312.6%16.5%estimated ± 5.1 pp, low confidence
MiniMax M32.0%6.0%estimated ± 5.1 pp, low confidence
Mistral Large 426.8%30.5%estimated ± 5.1 pp, medium confidence
Muse Glimmer 30B0.5%4.5%estimated ± 5.1 pp, low confidence
Muse Spark 1.333.3%37.0%estimated ± 5.1 pp, medium confidence
Nemotron 3 Ultra0.5%4.5%estimated ± 5.1 pp, low confidence
Qwen3.8-27B5.6%9.5%estimated ± 5.1 pp, low confidence