benchgap
Calibration

AA-SciCode → SciCode

SciCode is estimated from AA-SciCode with a linear curve fitted on 20 models measured on both: y = 0.9678·x + 0.0131, R² = 0.89, cross-validated error 3.2 pp. It is used for 54 estimates.

Estimated modelAA-SciCodeSciCodeSource
Apodex 1.145.5%45.3%estimated ± 3.2 pp, high confidence
Apodex 1.1 Mini45.5%45.3%estimated ± 3.2 pp, high confidence
Celeris-121.6%22.2%estimated ± 3.2 pp, medium confidence
Claude Haiku 5.555.0%54.5%estimated ± 3.2 pp, high confidence
Claude Opus 5.566.9%66.1%estimated ± 3.2 pp, medium confidence
Claude Sonnet 5.561.0%60.3%estimated ± 3.2 pp, medium confidence
Command A+38.5%38.6%estimated ± 3.2 pp, high confidence
DeepSeek V335.8%36.0%estimated ± 3.2 pp, high confidence
DeepSeek V3 032439.0%39.1%estimated ± 3.2 pp, high confidence
DeepSeek V4.1 Flash51.9%51.5%estimated ± 3.2 pp, high confidence
Gemini 2.5 Pro46.3%46.1%estimated ± 3.2 pp, high confidence
Gemini 4 Argon61.8%61.1%estimated ± 3.2 pp, medium confidence
Gemma 3 27B23.3%23.9%estimated ± 3.2 pp, medium confidence
Gemma 4 26B A4B40.0%40.0%estimated ± 3.2 pp, high confidence
Gemma 4 31B45.5%45.3%estimated ± 3.2 pp, high confidence
Gemma 4 E4B24.4%24.9%estimated ± 3.2 pp, medium confidence
GPT-5.6 Luna53.6%53.2%estimated ± 3.2 pp, high confidence
GPT-6.1 Sol54.2%53.8%estimated ± 3.2 pp, high confidence
GPT-6 Astra56.5%56.0%estimated ± 3.2 pp, high confidence
GPT-6 Luna54.6%54.2%estimated ± 3.2 pp, high confidence
GPT-6 Sol57.6%57.1%estimated ± 3.2 pp, high confidence
GPT-OSS 120B34.0%34.2%estimated ± 3.2 pp, high confidence
GPT-OSS 20B38.9%39.0%estimated ± 3.2 pp, high confidence
Grok 4.757.4%56.9%estimated ± 3.2 pp, high confidence
Hy348.6%48.3%estimated ± 3.2 pp, high confidence
LFM2.5-2.6B14.4%15.2%estimated ± 3.2 pp, medium confidence
Ling 3.0 Flash VL44.2%44.1%estimated ± 3.2 pp, high confidence
Ling 3.0 Tiny24.2%24.7%estimated ± 3.2 pp, medium confidence
Ling 3.1 Flash54.1%53.7%estimated ± 3.2 pp, high confidence
Llama 4 Maverick31.7%32.0%estimated ± 3.2 pp, high confidence
Llama 4 Scout21.3%21.9%estimated ± 3.2 pp, medium confidence
MiMo-V2.6-Flash51.3%51.0%estimated ± 3.2 pp, high confidence
MiMo-V2.6-Pro60.9%60.3%estimated ± 3.2 pp, medium confidence
Mistral Large 336.6%36.7%estimated ± 3.2 pp, high confidence
Mistral Large 454.2%53.8%estimated ± 3.2 pp, high confidence
Mistral Medium 3.5 128B40.2%40.2%estimated ± 3.2 pp, high confidence
Mistral Small 438.8%38.9%estimated ± 3.2 pp, high confidence
Mistral Small 4 (Reasoning)38.8%38.9%estimated ± 3.2 pp, high confidence
Muse Spark 1.257.4%56.9%estimated ± 3.2 pp, high confidence
Muse Spark 1.358.8%58.2%estimated ± 3.2 pp, high confidence
Nemotron 3 Nano 30B30.6%30.9%estimated ± 3.2 pp, high confidence
Nemotron 3 Super 100B36.2%36.3%estimated ± 3.2 pp, high confidence
North Mini Code38.8%38.9%estimated ± 3.2 pp, high confidence
Quasar 438B48.1%47.9%estimated ± 3.2 pp, high confidence
Qwen3.5-122B-A10B39.7%39.7%estimated ± 3.2 pp, high confidence
Qwen3.6-27B42.8%42.7%estimated ± 3.2 pp, high confidence
Qwen3.6-35B-A3B36.6%36.7%estimated ± 3.2 pp, high confidence
Qwen3.8-Flash-Next50.6%50.3%estimated ± 3.2 pp, high confidence
Qwen3.8 Max Preview52.1%51.7%estimated ± 3.2 pp, high confidence
Solar Pro 325.5%26.0%estimated ± 3.2 pp, high confidence
Solar Pro 444.6%44.5%estimated ± 3.2 pp, high confidence
Step 3.7 Flash43.9%43.8%estimated ± 3.2 pp, high confidence
Trinity-Large-Preview40.6%40.6%estimated ± 3.2 pp, high confidence
Trinity-Large-Thinking40.6%40.6%estimated ± 3.2 pp, high confidence