benchgap
Calibration

FrontierCode 1.1 Main → DeepSWE

DeepSWE is estimated from FrontierCode 1.1 Main with a Hill curve fitted on 6 models measured on both: y = 0.0000 + (0.7584 − 0.0000)·x^6.00 / (0.31281^6.00 + x^6.00), R² = 0.56, cross-validated error 3.4 pp. It is used for 10 estimates.

Estimated modelFrontierCode 1.1 MainDeepSWESource
Claude Fable 553.5%72.9%estimated ± 3.4 pp, medium confidence
Claude Haiku 5.546.4%69.3%estimated ± 3.4 pp, medium confidence
Claude Opus 4.626.9%21.8%estimated ± 3.4 pp, low confidence
Claude Opus 4.738.5%58.9%estimated ± 3.4 pp, low confidence
Claude Opus 4.846.5%69.4%estimated ± 3.4 pp, medium confidence
Claude Sonnet 4.624.3%13.7%estimated ± 3.4 pp, low confidence
Claude Sonnet 542.7%65.7%estimated ± 3.4 pp, low confidence
GPT-5.4 mini27.0%22.2%estimated ± 3.4 pp, low confidence
GPT-5.543.0%66.0%estimated ± 3.4 pp, low confidence
SWE-1.742.3%65.2%estimated ± 3.4 pp, low confidence