Calibration
Vals SWE-bench → DeepSWE
DeepSWE is estimated from Vals SWE-bench with a offset logistic curve fitted on 18 models measured on both: y = 0.5632 + (0.6782 − 0.5632) / (1 + exp(−200.00·(x − 0.9175))), R² = 0.43, cross-validated error 7.2 pp. It is used for 12 estimates.
| Estimated model | Vals SWE-bench | DeepSWE | Source |
|---|---|---|---|
| Claude Haiku 4.5 | 66.6% | 56.3% | estimated ± 7.2 pp, low confidence |
| Composer 2.5 | 79.6% | 56.3% | estimated ± 7.2 pp, low confidence |
| Gemini 3.1 Flash-Lite | 62.8% | 56.3% | estimated ± 7.2 pp, low confidence |
| Gemini 3 Flash | 75.0% | 56.3% | estimated ± 7.2 pp, low confidence |
| GPT-5.2-Codex | 72.4% | 56.3% | estimated ± 7.2 pp, low confidence |
| GPT-5.3 Codex | 78.0% | 56.3% | estimated ± 7.2 pp, low confidence |
| Grok 4.20 | 72.2% | 56.3% | estimated ± 7.2 pp, low confidence |
| Laguna M.1 | 57.6% | 56.3% | estimated ± 7.2 pp, low confidence |
| Laguna XS.2 | 55.2% | 56.3% | estimated ± 7.2 pp, low confidence |
| MiMo-V2.5 | 71.0% | 56.3% | estimated ± 7.2 pp, low confidence |
| Qwen3.5 Flash | 64.4% | 56.3% | estimated ± 7.2 pp, low confidence |
| Qwen 3.6 Max (preview) | 72.8% | 56.3% | estimated ± 7.2 pp, low confidence |