benchgap
Calibration

AA-SciCode → FrontierSWE v2

FrontierSWE v2 is estimated from AA-SciCode with a linear curve fitted on 17 models measured on both: y = 3.0836·x + -1.4255, R² = 0.49, cross-validated error 14.6 pp. It is used for 48 estimates.

Estimated modelAA-SciCodeFrontierSWE v2Source
A.X K241.0%0.0%estimated ± 14.6 pp, low confidence
Apodex 1.145.5%0.0%estimated ± 14.6 pp, low confidence
Apodex 1.1 Mini45.5%0.0%estimated ± 14.6 pp, low confidence
Celeris-121.6%0.0%estimated ± 14.6 pp, low confidence
Claude Haiku 5.555.0%27.1%estimated ± 14.6 pp, low confidence
Command A+38.5%0.0%estimated ± 14.6 pp, low confidence
DeepSeek V335.8%0.0%estimated ± 14.6 pp, low confidence
DeepSeek V3 032439.0%0.0%estimated ± 14.6 pp, low confidence
DeepSeek V4.1 Flash51.9%17.5%estimated ± 14.6 pp, low confidence
Gemma 3 27B23.3%0.0%estimated ± 14.6 pp, low confidence
Gemma 4 26B A4B40.0%0.0%estimated ± 14.6 pp, low confidence
Gemma 4 31B45.5%0.0%estimated ± 14.6 pp, low confidence
Gemma 4 E4B24.4%0.0%estimated ± 14.6 pp, low confidence
GPT-6.1 Sol54.2%24.6%estimated ± 14.6 pp, low confidence
GPT-6 Luna54.6%25.8%estimated ± 14.6 pp, low confidence
GPT-6 Sol57.6%35.1%estimated ± 14.6 pp, low confidence
GPT-OSS 120B34.0%0.0%estimated ± 14.6 pp, low confidence
GPT-OSS 20B38.9%0.0%estimated ± 14.6 pp, low confidence
Granite 4.2 3B25.3%0.0%estimated ± 14.6 pp, low confidence
Hy348.6%7.3%estimated ± 14.6 pp, low confidence
Hy3 Preview48.6%7.3%estimated ± 14.6 pp, low confidence
K-EXAONE 2.042.0%0.0%estimated ± 14.6 pp, low confidence
LFM2.5-2.6B14.4%0.0%estimated ± 14.6 pp, low confidence
Ling 3.0 Flash FP842.0%0.0%estimated ± 14.6 pp, low confidence
Ling 3.0 Flash VL44.2%0.0%estimated ± 14.6 pp, low confidence
Ling 3.0 Tiny24.2%0.0%estimated ± 14.6 pp, low confidence
Ling 3.1 Flash54.1%24.3%estimated ± 14.6 pp, low confidence
Llama 4 Maverick31.7%0.0%estimated ± 14.6 pp, low confidence
Llama 4 Scout21.3%0.0%estimated ± 14.6 pp, low confidence
Mercury 2.538.5%0.0%estimated ± 14.6 pp, low confidence
MiMo-V2.6-Flash51.3%15.6%estimated ± 14.6 pp, low confidence
MiMo-V2.6-Pro60.9%45.2%estimated ± 14.6 pp, low confidence
Mistral Large 336.6%0.0%estimated ± 14.6 pp, low confidence
Mistral Large 454.2%24.6%estimated ± 14.6 pp, low confidence
Mistral Small 438.8%0.0%estimated ± 14.6 pp, low confidence
Mistral Small 4 (Reasoning)38.8%0.0%estimated ± 14.6 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP432.1%0.0%estimated ± 14.6 pp, low confidence
Nemotron 3 Nano 30B30.6%0.0%estimated ± 14.6 pp, low confidence
Nemotron 3 Super 100B36.2%0.0%estimated ± 14.6 pp, low confidence
North Mini Code38.8%0.0%estimated ± 14.6 pp, low confidence
Quasar 438B48.1%5.8%estimated ± 14.6 pp, low confidence
Qwen3.5-122B-A10B39.7%0.0%estimated ± 14.6 pp, low confidence
Qwen3.8 Max Preview52.1%18.1%estimated ± 14.6 pp, low confidence
Solar Pro 325.5%0.0%estimated ± 14.6 pp, low confidence
Solar Pro 444.6%0.0%estimated ± 14.6 pp, low confidence
Step 5 Preview58.9%39.1%estimated ± 14.6 pp, low confidence
Trinity-Large-Preview40.6%0.0%estimated ± 14.6 pp, low confidence
Trinity-Large-Thinking40.6%0.0%estimated ± 14.6 pp, low confidence