benchgap
Calibration

SWE-bench Pro → AA-SciCode

AA-SciCode is estimated from SWE-bench Pro with a linear curve fitted on 37 models measured on both: y = 0.5237·x + 0.1893, R² = 0.81, cross-validated error 3.8 pp. It is used for 14 estimates.

Estimated modelSWE-bench ProAA-SciCodeSource
Atria Dawn Preview59.6%50.1%estimated ± 3.8 pp, high confidence
Claude Mythos 580.3%61.0%estimated ± 3.8 pp, high confidence
Claude Opus 4.653.4%46.9%estimated ± 3.8 pp, high confidence
GLM-555.1%47.8%estimated ± 3.8 pp, high confidence
GPT-5.255.6%48.1%estimated ± 3.8 pp, high confidence
GPT-5.3 Codex56.8%48.7%estimated ± 3.8 pp, high confidence
Grok 4.2051.8%46.1%estimated ± 3.8 pp, high confidence
Laguna M.149.2%44.7%estimated ± 3.8 pp, high confidence
Laguna S 2.159.4%50.0%estimated ± 3.8 pp, high confidence
Laguna XS.246.3%43.2%estimated ± 3.8 pp, high confidence
Laguna XS 2.147.6%43.9%estimated ± 3.8 pp, high confidence
LLaDA2.2-flash30.1%34.7%estimated ± 3.8 pp, high confidence
LongCat-Flash-Lite-Sparse40.6%40.2%estimated ± 3.8 pp, high confidence
MiMo-V2.556.1%48.3%estimated ± 3.8 pp, high confidence