benchgap
Calibration

AA-SciCode → Vals SWE-bench

Vals SWE-bench is estimated from AA-SciCode with a linear curve fitted on 42 models measured on both: y = 1.2595·x + 0.1627, R² = 0.47, cross-validated error 7.6 pp. It is used for 11 estimates.

Estimated modelAA-SciCodeVals SWE-benchSource
DeepSeek V3 032439.0%65.4%estimated ± 7.6 pp, low confidence
GPT-6.1 Sol54.2%84.5%estimated ± 7.6 pp, low confidence
GPT-6 Luna54.6%85.0%estimated ± 7.6 pp, low confidence
GPT-6 Sol57.6%88.8%estimated ± 7.6 pp, low confidence
Ling 3.0 Flash VL44.2%71.9%estimated ± 7.6 pp, low confidence
Ling 3.0 Tiny24.2%46.7%estimated ± 7.6 pp, low confidence
Ling 3.1 Flash54.1%84.4%estimated ± 7.6 pp, low confidence
MiMo-V2.6-Flash51.3%80.9%estimated ± 7.6 pp, low confidence
MiMo-V2.6-Pro60.9%93.0%estimated ± 7.6 pp, low confidence
North Mini Code38.8%65.1%estimated ± 7.6 pp, low confidence
Solar Pro 325.5%48.4%estimated ± 7.6 pp, low confidence