Calibration
Vals SWE-bench → FrontierCode 1.1 Extended
FrontierCode 1.1 Extended is estimated from Vals SWE-bench with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.95815^6.00 + x^6.00), R² = 0.69, cross-validated error 2.4 pp. It is used for 43 estimates.
| Estimated model | Vals SWE-bench | FrontierCode 1.1 Extended | Source |
|---|---|---|---|
| Claude Haiku 4.5 | 66.6% | 12.2% | estimated ± 2.4 pp, low confidence |
| Claude Opus 4.7 | 82.0% | 33.8% | estimated ± 2.4 pp, low confidence |
| Claude Sonnet 4.6 | 77.4% | 26.1% | estimated ± 2.4 pp, low confidence |
| DeepSeek V4 Flash 0731 | 88.8% | 46.5% | estimated ± 2.4 pp, low confidence |
| DeepSeek V4 Pro 0813 | 96.4% | 61.1% | estimated ± 2.4 pp, medium confidence |
| Gemini 2.5 Pro | 54.4% | 3.9% | estimated ± 2.4 pp, low confidence |
| Gemini 3.1 Flash-Lite | 62.8% | 8.8% | estimated ± 2.4 pp, low confidence |
| Gemini 3.1 Pro | 78.8% | 28.4% | estimated ± 2.4 pp, low confidence |
| Gemini 3.5 Flash-Lite | 75.0% | 22.4% | estimated ± 2.4 pp, low confidence |
| Gemini 3.7 Flash | 80.8% | 31.7% | estimated ± 2.4 pp, low confidence |
| Gemini 3 Flash | 75.0% | 22.4% | estimated ± 2.4 pp, low confidence |
| GLM-4.7 | 69.4% | 15.1% | estimated ± 2.4 pp, low confidence |
| GLM-5.1 | 76.4% | 24.5% | estimated ± 2.4 pp, low confidence |
| GLM-5.3 | 95.4% | 59.2% | estimated ± 2.4 pp, medium confidence |
| GLM-5.3-Flash | 92.0% | 52.7% | estimated ± 2.4 pp, low confidence |
| GPT-5.2-Codex | 72.4% | 18.8% | estimated ± 2.4 pp, low confidence |
| GPT-5.3 Codex | 78.0% | 27.1% | estimated ± 2.4 pp, low confidence |
| GPT-5.4 mini | 73.0% | 19.6% | estimated ± 2.4 pp, low confidence |
| GPT-5.4 nano | 69.8% | 15.6% | estimated ± 2.4 pp, low confidence |
| Grok 4.20 | 72.2% | 18.6% | estimated ± 2.4 pp, low confidence |
| Grok 4.3 | 71.4% | 17.5% | estimated ± 2.4 pp, low confidence |
| Inkling | 77.6% | 26.4% | estimated ± 2.4 pp, low confidence |
| Inkling-Small | 82.2% | 34.2% | estimated ± 2.4 pp, low confidence |
| Kimi K2.6 | 76.2% | 24.2% | estimated ± 2.4 pp, low confidence |
| Laguna M.1 | 57.6% | 5.4% | estimated ± 2.4 pp, low confidence |
| Laguna XS.2 | 55.2% | 4.2% | estimated ± 2.4 pp, low confidence |
| Ling 3.0 Flash | 65.2% | 10.8% | estimated ± 2.4 pp, low confidence |
| MiMo-V2.5 | 71.0% | 17.0% | estimated ± 2.4 pp, low confidence |
| MiMo-V2.5-Pro | 74.0% | 21.0% | estimated ± 2.4 pp, low confidence |
| MiniMax M2.7 | 73.8% | 20.7% | estimated ± 2.4 pp, low confidence |
| MiniMax M3 | 75.0% | 22.4% | estimated ± 2.4 pp, low confidence |
| Mistral Medium 3.5 128B | 66.4% | 12.0% | estimated ± 2.4 pp, low confidence |
| Muse Spark | 74.4% | 21.6% | estimated ± 2.4 pp, low confidence |
| Muse Spark 1.1 | 82.0% | 33.8% | estimated ± 2.4 pp, low confidence |
| Muse Spark 1.2 | 86.6% | 42.3% | estimated ± 2.4 pp, low confidence |
| Nemotron 3 Ultra | 69.0% | 14.7% | estimated ± 2.4 pp, low confidence |
| Qwen3.5 Flash | 64.4% | 10.1% | estimated ± 2.4 pp, low confidence |
| Qwen3.6-27B | 70.0% | 15.8% | estimated ± 2.4 pp, low confidence |
| Qwen 3.6 Max (preview) | 72.8% | 19.4% | estimated ± 2.4 pp, low confidence |
| Qwen3.6 Plus | 73.4% | 20.2% | estimated ± 2.4 pp, low confidence |
| Qwen3.7 Max | 68.8% | 14.5% | estimated ± 2.4 pp, low confidence |
| Qwen3.8-27B | 86.0% | 41.2% | estimated ± 2.4 pp, low confidence |
| Qwen3.8 Max | 85.6% | 40.4% | estimated ± 2.4 pp, low confidence |