benchgap
Calibration

AA-SciCode → SWE-bench Pro

SWE-bench Pro is estimated from AA-SciCode with a linear curve fitted on 37 models measured on both: y = 1.5480·x + -0.1818, R² = 0.81, cross-validated error 6.8 pp. It is used for 15 estimates.

Estimated modelAA-SciCodeSWE-bench ProSource
Claude Haiku 5.555.0%67.0%estimated ± 6.8 pp, medium confidence
DeepSeek V3 032439.0%42.2%estimated ± 6.8 pp, medium confidence
GPT-6.1 Sol54.2%65.7%estimated ± 6.8 pp, medium confidence
GPT-6 Luna54.6%66.3%estimated ± 6.8 pp, medium confidence
GPT-6 Sol57.6%71.0%estimated ± 6.8 pp, medium confidence
Ling 3.0 Flash VL44.2%50.2%estimated ± 6.8 pp, medium confidence
Ling 3.0 Tiny24.2%19.3%estimated ± 6.8 pp, low confidence
Ling 3.1 Flash54.1%65.6%estimated ± 6.8 pp, medium confidence
Mercury 2.538.5%41.4%estimated ± 6.8 pp, medium confidence
MiMo-V2.6-Flash51.3%61.2%estimated ± 6.8 pp, medium confidence
MiMo-V2.6-Pro60.9%76.1%estimated ± 6.8 pp, medium confidence
Mistral Large 454.2%65.7%estimated ± 6.8 pp, medium confidence
North Mini Code38.8%41.9%estimated ± 6.8 pp, medium confidence
Solar Pro 325.5%21.3%estimated ± 6.8 pp, low confidence
Step 5 Preview58.9%73.0%estimated ± 6.8 pp, medium confidence