Calibration
FrontierCode 1.1 Main → DeepSWE
DeepSWE is estimated from FrontierCode 1.1 Main with a Hill curve fitted on 6 models measured on both: y = 0.0000 + (0.7584 − 0.0000)·x^6.00 / (0.31281^6.00 + x^6.00), R² = 0.56, cross-validated error 3.4 pp. It is used for 10 estimates.
| Estimated model | FrontierCode 1.1 Main | DeepSWE | Source |
|---|---|---|---|
| Claude Fable 5 | 53.5% | 72.9% | estimated ± 3.4 pp, medium confidence |
| Claude Haiku 5.5 | 46.4% | 69.3% | estimated ± 3.4 pp, medium confidence |
| Claude Opus 4.6 | 26.9% | 21.8% | estimated ± 3.4 pp, low confidence |
| Claude Opus 4.7 | 38.5% | 58.9% | estimated ± 3.4 pp, low confidence |
| Claude Opus 4.8 | 46.5% | 69.4% | estimated ± 3.4 pp, medium confidence |
| Claude Sonnet 4.6 | 24.3% | 13.7% | estimated ± 3.4 pp, low confidence |
| Claude Sonnet 5 | 42.7% | 65.7% | estimated ± 3.4 pp, low confidence |
| GPT-5.4 mini | 27.0% | 22.2% | estimated ± 3.4 pp, low confidence |
| GPT-5.5 | 43.0% | 66.0% | estimated ± 3.4 pp, low confidence |
| SWE-1.7 | 42.3% | 65.2% | estimated ± 3.4 pp, low confidence |