benchgap
Calibration

Vals SWE-bench → DeepSWE

DeepSWE is estimated from Vals SWE-bench with a offset logistic curve fitted on 18 models measured on both: y = 0.5632 + (0.6782 − 0.5632) / (1 + exp(−200.00·(x − 0.9175))), R² = 0.43, cross-validated error 7.2 pp. It is used for 12 estimates.

Estimated modelVals SWE-benchDeepSWESource
Claude Haiku 4.566.6%56.3%estimated ± 7.2 pp, low confidence
Composer 2.579.6%56.3%estimated ± 7.2 pp, low confidence
Gemini 3.1 Flash-Lite62.8%56.3%estimated ± 7.2 pp, low confidence
Gemini 3 Flash75.0%56.3%estimated ± 7.2 pp, low confidence
GPT-5.2-Codex72.4%56.3%estimated ± 7.2 pp, low confidence
GPT-5.3 Codex78.0%56.3%estimated ± 7.2 pp, low confidence
Grok 4.2072.2%56.3%estimated ± 7.2 pp, low confidence
Laguna M.157.6%56.3%estimated ± 7.2 pp, low confidence
Laguna XS.255.2%56.3%estimated ± 7.2 pp, low confidence
MiMo-V2.571.0%56.3%estimated ± 7.2 pp, low confidence
Qwen3.5 Flash64.4%56.3%estimated ± 7.2 pp, low confidence
Qwen 3.6 Max (preview)72.8%56.3%estimated ± 7.2 pp, low confidence