benchgap
Calibration

Vals LiveCodeBench → SWE-bench Pro

SWE-bench Pro is estimated from Vals LiveCodeBench with a offset logistic curve fitted on 32 models measured on both: y = 0.5697 + (0.8227 − 0.5697) / (1 + exp(−200.00·(x − 0.8833))), R² = 0.69, cross-validated error 4.8 pp. It is used for 11 estimates.

Estimated modelVals LiveCodeBenchSWE-bench ProSource
Gemini 3.7 Flash88.7%74.1%estimated ± 4.8 pp, high confidence
GPT-5.2-Codex88.0%65.6%estimated ± 4.8 pp, high confidence
Gemini 3 Flash85.6%57.1%estimated ± 4.8 pp, high confidence
GPT-5.1-Codex85.6%57.1%estimated ± 4.8 pp, high confidence
Claude Opus 4.785.1%57.0%estimated ± 4.8 pp, high confidence
Grok 4.384.5%57.0%estimated ± 4.8 pp, high confidence
GPT-5.1-Codex-Max83.6%57.0%estimated ± 4.8 pp, high confidence
Qwen3.5 Flash83.3%57.0%estimated ± 4.8 pp, high confidence
GLM-4.681.0%57.0%estimated ± 4.8 pp, high confidence
Gemini 3.1 Flash-Lite80.1%57.0%estimated ± 4.8 pp, high confidence
GLM-4.567.4%57.0%estimated ± 4.8 pp, medium confidence