benchgap
Calibration

SWE-bench Pro → LiveCodeBench v6

LiveCodeBench v6 is estimated from SWE-bench Pro with a linear curve fitted on 15 models measured on both: y = 0.4408·x + 0.6312, R² = 0.93, cross-validated error 2.1 pp. It is used for 24 estimates.

Estimated modelSWE-bench ProLiveCodeBench v6Source
Atria Dawn Preview59.6%89.4%estimated ± 2.1 pp, high confidence
Claude Fable 5.181.2%98.9%estimated ± 2.1 pp, medium confidence
Claude Opus 5.589.9%100.0%estimated ± 2.1 pp, medium confidence
Claude Sonnet 5.581.3%99.0%estimated ± 2.1 pp, medium confidence
Gemini 3.5 Flash55.1%87.4%estimated ± 2.1 pp, high confidence
Gemini 3.5 Flash-Lite54.2%87.0%estimated ± 2.1 pp, high confidence
GLM-5.158.4%88.9%estimated ± 2.1 pp, high confidence
GLM-5.262.1%90.5%estimated ± 2.1 pp, high confidence
GPT-5.457.7%88.6%estimated ± 2.1 pp, high confidence
GPT-5.558.6%89.0%estimated ± 2.1 pp, high confidence
GPT-5.6 Luna62.7%90.8%estimated ± 2.1 pp, high confidence
GPT-5.6 Sol64.6%91.6%estimated ± 2.1 pp, high confidence
GPT-5.6 Terra63.4%91.1%estimated ± 2.1 pp, high confidence
Grok 4.564.7%91.6%estimated ± 2.1 pp, high confidence
Hy4 preview65.7%92.1%estimated ± 2.1 pp, high confidence
Laguna S 2.159.4%89.3%estimated ± 2.1 pp, high confidence
Ling 3.0 Flash56.6%88.1%estimated ± 2.1 pp, high confidence
MiMo-V2.556.1%87.9%estimated ± 2.1 pp, high confidence
MiMo-V2.5-Pro57.2%88.3%estimated ± 2.1 pp, high confidence
MiniMax M2.756.2%87.9%estimated ± 2.1 pp, high confidence
Muse Spark 1.161.5%90.2%estimated ± 2.1 pp, high confidence
Qwen 3.6 Max (preview)57.3%88.4%estimated ± 2.1 pp, high confidence
Qwen3.8 Max67.7%93.0%estimated ± 2.1 pp, high confidence
Step 3.7 Flash56.3%87.9%estimated ± 2.1 pp, high confidence