benchgap
Calibration

SWE-bench Pro → LiveCodeBench Pro

LiveCodeBench Pro is estimated from SWE-bench Pro with a Michaelis–Menten curve fitted on 6 models measured on both: y = 1.8827·x / (0.74902 + x), R² = 0.60, cross-validated error 6.0 pp. It is used for 67 estimates.

Estimated modelSWE-bench ProLiveCodeBench ProSource
Atria Dawn Preview59.6%83.4%estimated ± 6.0 pp, low confidence
Claude Fable 580.0%97.2%estimated ± 6.0 pp, low confidence
Claude Fable 5.181.2%97.9%estimated ± 6.0 pp, low confidence
Claude Mythos 580.3%97.4%estimated ± 6.0 pp, low confidence
Claude Opus 4.557.1%81.4%estimated ± 6.0 pp, low confidence
Claude Opus 4.7 (Adaptive)64.3%87.0%estimated ± 6.0 pp, low confidence
Claude Opus 4.869.2%90.4%estimated ± 6.0 pp, low confidence
Claude Opus 579.2%96.8%estimated ± 6.0 pp, low confidence
Claude Opus 5.589.9%100.0%estimated ± 6.0 pp, low confidence
Claude Sonnet 563.2%86.2%estimated ± 6.0 pp, low confidence
Claude Sonnet 5.581.3%98.0%estimated ± 6.0 pp, low confidence
DeepSeek V4 Flash 073152.6%77.7%estimated ± 6.0 pp, low confidence
DeepSeek V4 Pro 081355.4%80.0%estimated ± 6.0 pp, low confidence
dots3-note Preview61.0%84.5%estimated ± 6.0 pp, low confidence
Gemini 3.5 Flash55.1%79.8%estimated ± 6.0 pp, low confidence
Gemini 3.5 Flash-Lite54.2%79.0%estimated ± 6.0 pp, low confidence
GLM-555.1%79.8%estimated ± 6.0 pp, low confidence
GLM-5.158.4%82.5%estimated ± 6.0 pp, low confidence
GLM-5.262.1%85.3%estimated ± 6.0 pp, low confidence
GPT-5.255.6%80.2%estimated ± 6.0 pp, low confidence
GPT-5.3 Codex56.8%81.2%estimated ± 6.0 pp, low confidence
GPT-5.558.6%82.6%estimated ± 6.0 pp, low confidence
GPT-5.6 Luna62.7%85.8%estimated ± 6.0 pp, low confidence
GPT-5.6 Sol64.6%87.2%estimated ± 6.0 pp, low confidence
GPT-5.6 Terra63.4%86.3%estimated ± 6.0 pp, low confidence
Granite 4.2 30B33.3%57.9%estimated ± 6.0 pp, low confidence
Granite 4.2 8B19.1%38.3%estimated ± 6.0 pp, low confidence
Grok 4.564.7%87.3%estimated ± 6.0 pp, low confidence
Hy4 preview65.7%88.0%estimated ± 6.0 pp, low confidence
Inkling54.3%79.1%estimated ± 6.0 pp, low confidence
Inkling-Small55.9%80.5%estimated ± 6.0 pp, low confidence
Kimi K2.658.6%82.6%estimated ± 6.0 pp, low confidence
Kimi K2.550.7%76.0%estimated ± 6.0 pp, low confidence
Laguna M.149.2%74.6%estimated ± 6.0 pp, low confidence
Laguna S 2.159.4%83.3%estimated ± 6.0 pp, low confidence
Laguna XS.246.3%71.9%estimated ± 6.0 pp, low confidence
Laguna XS 2.147.6%73.2%estimated ± 6.0 pp, low confidence
Ling 3.0 Flash56.6%81.0%estimated ± 6.0 pp, low confidence
LLaDA2.2-flash30.1%54.0%estimated ± 6.0 pp, low confidence
LongCat-Flash-Lite-Sparse40.6%66.2%estimated ± 6.0 pp, low confidence
MAI-Thinking-152.8%77.8%estimated ± 6.0 pp, low confidence
MiMo-V2.556.1%80.6%estimated ± 6.0 pp, low confidence
MiMo-V2.5-Pro57.2%81.5%estimated ± 6.0 pp, low confidence
MiniCPM5-2B14.4%30.4%estimated ± 6.0 pp, low confidence
MiniMax M2.756.2%80.7%estimated ± 6.0 pp, low confidence
MiniMax M359.0%83.0%estimated ± 6.0 pp, low confidence
Muse Glimmer 30B51.2%76.4%estimated ± 6.0 pp, low confidence
Muse Spark 1.161.5%84.9%estimated ± 6.0 pp, low confidence
Ornith-1.0-35B50.4%75.7%estimated ± 6.0 pp, low confidence
Ornith-1.0-397B62.2%85.4%estimated ± 6.0 pp, low confidence
Ornith-1.0-9B42.9%68.6%estimated ± 6.0 pp, low confidence
Ornith-1.5-35B-A3B59.6%83.4%estimated ± 6.0 pp, low confidence
Ornith-1.5-397B65.1%87.5%estimated ± 6.0 pp, low confidence
Ornith-1.5-9B47.5%73.1%estimated ± 6.0 pp, low confidence
Qwen3.5 397B50.9%76.2%estimated ± 6.0 pp, low confidence
Qwen3.6-27B53.5%78.4%estimated ± 6.0 pp, low confidence
Qwen3.6-35B-A3B49.5%74.9%estimated ± 6.0 pp, low confidence
Qwen 3.6 Max (preview)57.3%81.6%estimated ± 6.0 pp, low confidence
Qwen3.6 Plus56.6%81.0%estimated ± 6.0 pp, low confidence
Qwen3.7 Max60.6%84.2%estimated ± 6.0 pp, low confidence
Qwen3.7 Plus57.6%81.8%estimated ± 6.0 pp, low confidence
Qwen3.8-27B61.7%85.0%estimated ± 6.0 pp, low confidence
Qwen3.8-Flash-Next62.5%85.6%estimated ± 6.0 pp, low confidence
Qwen3.8 Max67.7%89.4%estimated ± 6.0 pp, low confidence
Qwen3.8-Omni-Flash63.3%86.2%estimated ± 6.0 pp, low confidence
Beam65.5%87.8%estimated ± 6.0 pp, low confidence
Step 3.7 Flash56.3%80.8%estimated ± 6.0 pp, low confidence