benchgap
Calibration

SWE-bench Pro → Vals LiveCodeBench

Vals LiveCodeBench is estimated from SWE-bench Pro with a Hill curve fitted on 32 models measured on both: y = 0.0000 + (0.8931 − 0.0000)·x^6.00 / (0.36733^6.00 + x^6.00), R² = 0.40, cross-validated error 4.7 pp. It is used for 25 estimates.

Estimated modelSWE-bench ProVals LiveCodeBenchSource
Atria Dawn Preview59.6%84.7%estimated ± 4.7 pp, medium confidence
Claude Mythos 580.3%88.5%estimated ± 4.7 pp, medium confidence
Claude Opus 4.557.1%83.4%estimated ± 4.7 pp, medium confidence
Claude Opus 4.7 (Adaptive)64.3%86.3%estimated ± 4.7 pp, medium confidence
dots3-note Preview61.0%85.2%estimated ± 4.7 pp, medium confidence
GPT-5.255.6%82.5%estimated ± 4.7 pp, medium confidence
Hy4 preview65.7%86.7%estimated ± 4.7 pp, medium confidence
Laguna S 2.159.4%84.6%estimated ± 4.7 pp, medium confidence
Laguna XS 2.147.6%73.7%estimated ± 4.7 pp, medium confidence
LLaDA2.2-flash30.1%20.8%estimated ± 4.7 pp, low confidence
LongCat-Flash-Lite-Sparse40.6%57.8%estimated ± 4.7 pp, low confidence
MAI-Thinking-152.8%80.2%estimated ± 4.7 pp, medium confidence
Muse Spark52.4%79.8%estimated ± 4.7 pp, medium confidence
Ornith-1.0-35B50.4%77.7%estimated ± 4.7 pp, medium confidence
Ornith-1.0-397B62.2%85.7%estimated ± 4.7 pp, medium confidence
Ornith-1.0-9B42.9%64.1%estimated ± 4.7 pp, low confidence
Ornith-1.5-35B-A3B59.6%84.7%estimated ± 4.7 pp, medium confidence
Ornith-1.5-397B65.1%86.5%estimated ± 4.7 pp, medium confidence
Ornith-1.5-9B47.5%73.6%estimated ± 4.7 pp, medium confidence
Qwen3.5 397B50.9%78.3%estimated ± 4.7 pp, medium confidence
Qwen3.6-27B53.5%80.8%estimated ± 4.7 pp, medium confidence
Qwen3.6-35B-A3B49.5%76.5%estimated ± 4.7 pp, medium confidence
Qwen3.8-Flash-Next62.5%85.8%estimated ± 4.7 pp, medium confidence
Qwen3.8-Omni-Flash63.3%86.0%estimated ± 4.7 pp, medium confidence
Step 3.7 Flash56.3%82.9%estimated ± 4.7 pp, medium confidence