Calibration
Vals SWE-bench → FrontierSWE v2
FrontierSWE v2 is estimated from Vals SWE-bench with a inverse Michaelis–Menten curve fitted on 12 models measured on both: y = 0.03090·(x − 0.0000) / (0.0000 + 1.0390 − x), R² = 0.67, cross-validated error 9.5 pp. It is used for 17 estimates.
| Estimated model | Vals SWE-bench | FrontierSWE v2 | Source |
|---|---|---|---|
| Claude Haiku 4.5 | 66.6% | 5.5% | estimated ± 9.5 pp, low confidence |
| Claude Opus 4.7 | 82.0% | 11.6% | estimated ± 9.5 pp, medium confidence |
| Claude Sonnet 4.6 | 77.4% | 9.0% | estimated ± 9.5 pp, low confidence |
| Gemini 2.5 Pro | 54.4% | 3.4% | estimated ± 9.5 pp, low confidence |
| Gemini 3.1 Flash-Lite | 62.8% | 4.7% | estimated ± 9.5 pp, low confidence |
| Gemini 3.1 Pro | 78.8% | 9.7% | estimated ± 9.5 pp, medium confidence |
| Gemini 3.6 Flash | 79.6% | 10.1% | estimated ± 9.5 pp, medium confidence |
| Gemini 3 Flash | 75.0% | 8.0% | estimated ± 9.5 pp, low confidence |
| GLM-4.7 | 69.4% | 6.2% | estimated ± 9.5 pp, low confidence |
| GPT-5.2-Codex | 72.4% | 7.1% | estimated ± 9.5 pp, low confidence |
| GPT-5.4 mini | 73.0% | 7.3% | estimated ± 9.5 pp, low confidence |
| GPT-5.4 nano | 69.8% | 6.3% | estimated ± 9.5 pp, low confidence |
| Grok 4.3 | 71.4% | 6.8% | estimated ± 9.5 pp, low confidence |
| Kimi K2.7 Code | 78.2% | 9.4% | estimated ± 9.5 pp, medium confidence |
| Mistral Medium 3.5 128B | 66.4% | 5.5% | estimated ± 9.5 pp, low confidence |
| Nemotron 3 Ultra | 69.0% | 6.1% | estimated ± 9.5 pp, low confidence |
| Qwen3.5 Flash | 64.4% | 5.0% | estimated ± 9.5 pp, low confidence |