benchgap
Calibration

Terminal-Bench 2.1 → AA Terminal-Bench 2.1

AA Terminal-Bench 2.1 is estimated from Terminal-Bench 2.1 with a offset logistic curve fitted on 10 models measured on both: y = 0.5207 + (0.8493 − 0.5207) / (1 + exp(−37.51·(x − 0.6788))), R² = 0.99, cross-validated error 3.5 pp. It is used for 51 estimates.

Estimated modelTerminal-Bench 2.1AA Terminal-Bench 2.1Source
A.X K236.0%52.1%estimated ± 3.5 pp, medium confidence
Apodex 1.170.8%76.7%estimated ± 3.5 pp, high confidence
Atria Dawn Preview78.3%84.3%estimated ± 3.5 pp, high confidence
Claude Fable 584.3%84.9%estimated ± 3.5 pp, high confidence
Claude Mythos 588.0%84.9%estimated ± 3.5 pp, high confidence
Claude Opus 4.874.6%82.5%estimated ± 3.5 pp, high confidence
Claude Sonnet 580.4%84.6%estimated ± 3.5 pp, high confidence
DeepSeek V4 Flash 073182.7%84.8%estimated ± 3.5 pp, high confidence
DeepSeek V4 Pro 081387.9%84.9%estimated ± 3.5 pp, high confidence
dots3-note Preview75.1%82.9%estimated ± 3.5 pp, high confidence
Ember-182.0%84.8%estimated ± 3.5 pp, high confidence
Gemini 3.5 Flash76.2%83.5%estimated ± 3.5 pp, high confidence
Gemini 3.5 Flash-Lite54.0%52.3%estimated ± 3.5 pp, high confidence
Gemini 3.7 Flash85.8%84.9%estimated ± 3.5 pp, high confidence
GLM-5.281.0%84.7%estimated ± 3.5 pp, high confidence
GPT-5.6 Luna84.7%84.9%estimated ± 3.5 pp, high confidence
GPT-5.6 Sol91.9%84.9%estimated ± 3.5 pp, medium confidence
GPT-5.6 Terra87.4%84.9%estimated ± 3.5 pp, high confidence
Granite 4.2 30B29.2%52.1%estimated ± 3.5 pp, medium confidence
Granite 4.2 8B20.6%52.1%estimated ± 3.5 pp, medium confidence
Grok 4.583.3%84.8%estimated ± 3.5 pp, high confidence
Hy4 preview85.4%84.9%estimated ± 3.5 pp, high confidence
Inkling-Small64.7%59.7%estimated ± 3.5 pp, high confidence
K-EXAONE 2.043.8%52.1%estimated ± 3.5 pp, medium confidence
Laguna S 2.170.2%75.2%estimated ± 3.5 pp, high confidence
Ling 3.0 Flash57.0%52.6%estimated ± 3.5 pp, high confidence
MAI-Code-1.1-Flash62.9%56.5%estimated ± 3.5 pp, high confidence
MiMo-V2.6-Flash87.6%84.9%estimated ± 3.5 pp, high confidence
MiniCPM5-2B8.6%52.1%estimated ± 3.5 pp, medium confidence
Muse Spark 1.180.0%84.6%estimated ± 3.5 pp, high confidence
Muse Spark 1.282.9%84.8%estimated ± 3.5 pp, high confidence
Nemotron 3.5 Lightning 30B A3B NVFP423.5%52.1%estimated ± 3.5 pp, medium confidence
Ornith-1.0-35B64.2%58.7%estimated ± 3.5 pp, high confidence
Ornith-1.0-397B77.5%84.1%estimated ± 3.5 pp, high confidence
Ornith-1.0-9B43.1%52.1%estimated ± 3.5 pp, medium confidence
Ornith-1.5-35B-A3B67.8%68.3%estimated ± 3.5 pp, high confidence
Ornith-1.5-397B86.1%84.9%estimated ± 3.5 pp, high confidence
Ornith-1.5-9B46.2%52.1%estimated ± 3.5 pp, medium confidence
Pokee-Isaac 28B65.1%60.6%estimated ± 3.5 pp, high confidence
Quasar 438B69.3%72.8%estimated ± 3.5 pp, high confidence
Qwen3.8 Max86.6%84.9%estimated ± 3.5 pp, high confidence
Beam80.1%84.6%estimated ± 3.5 pp, high confidence
Sakana Fugu80.2%84.6%estimated ± 3.5 pp, high confidence
Sakana Fugu-Ultra82.1%84.8%estimated ± 3.5 pp, high confidence
Seed 2.1 Pro71.0%77.2%estimated ± 3.5 pp, high confidence
Seed 2.1 Turbo67.6%67.6%estimated ± 3.5 pp, high confidence
Solar Pro 457.0%52.6%estimated ± 3.5 pp, high confidence
Step 3.7 Flash59.5%53.4%estimated ± 3.5 pp, high confidence
SWE-1.781.5%84.7%estimated ± 3.5 pp, high confidence
SWE-292.8%84.9%estimated ± 3.5 pp, medium confidence
Ternary Bonsai 2 27B52.8%52.2%estimated ± 3.5 pp, high confidence