benchgap
Calibration

AA-SciCode → LiveCodeBench v6

LiveCodeBench v6 is estimated from AA-SciCode with a linear curve fitted on 10 models measured on both: y = 0.9026·x + 0.4623, R² = 0.93, cross-validated error 3.3 pp. It is used for 21 estimates.

Estimated modelAA-SciCodeLiveCodeBench v6Source
Claude Haiku 5.555.0%95.9%estimated ± 3.3 pp, medium confidence
DeepSeek V3 032439.0%81.4%estimated ± 3.3 pp, high confidence
DeepSeek V4.1 Flash51.9%93.1%estimated ± 3.3 pp, medium confidence
Gemini 4 Argon61.8%100.0%estimated ± 3.3 pp, medium confidence
GLM-5.3-Flash51.6%92.8%estimated ± 3.3 pp, medium confidence
GPT-6.1 Sol54.2%95.2%estimated ± 3.3 pp, medium confidence
GPT-6 Luna54.6%95.5%estimated ± 3.3 pp, medium confidence
GPT-6 Sol57.6%98.2%estimated ± 3.3 pp, medium confidence
Grok 4.757.4%98.0%estimated ± 3.3 pp, medium confidence
K-EXAONE 2.042.0%84.1%estimated ± 3.3 pp, high confidence
Ling 3.0 Flash VL44.2%86.1%estimated ± 3.3 pp, high confidence
Ling 3.0 Tiny24.2%68.1%estimated ± 3.3 pp, high confidence
Ling 3.1 Flash54.1%95.1%estimated ± 3.3 pp, medium confidence
Mercury 2.538.5%81.0%estimated ± 3.3 pp, high confidence
MiMo-V2.6-Flash51.3%92.5%estimated ± 3.3 pp, high confidence
MiMo-V2.6-Pro60.9%100.0%estimated ± 3.3 pp, medium confidence
Mistral Large 454.2%95.2%estimated ± 3.3 pp, medium confidence
North Mini Code38.8%81.3%estimated ± 3.3 pp, high confidence
Solar Pro 325.5%69.2%estimated ± 3.3 pp, high confidence
Solar Pro 444.6%86.5%estimated ± 3.3 pp, high confidence
Step 5 Preview58.9%99.4%estimated ± 3.3 pp, medium confidence