benchgap
Calibration

Vals SWE-bench → FrontierSWE v2

FrontierSWE v2 is estimated from Vals SWE-bench with a inverse Michaelis–Menten curve fitted on 12 models measured on both: y = 0.03090·(x − 0.0000) / (0.0000 + 1.0390 − x), R² = 0.67, cross-validated error 9.5 pp. It is used for 17 estimates.

Estimated modelVals SWE-benchFrontierSWE v2Source
Claude Haiku 4.566.6%5.5%estimated ± 9.5 pp, low confidence
Claude Opus 4.782.0%11.6%estimated ± 9.5 pp, medium confidence
Claude Sonnet 4.677.4%9.0%estimated ± 9.5 pp, low confidence
Gemini 2.5 Pro54.4%3.4%estimated ± 9.5 pp, low confidence
Gemini 3.1 Flash-Lite62.8%4.7%estimated ± 9.5 pp, low confidence
Gemini 3.1 Pro78.8%9.7%estimated ± 9.5 pp, medium confidence
Gemini 3.6 Flash79.6%10.1%estimated ± 9.5 pp, medium confidence
Gemini 3 Flash75.0%8.0%estimated ± 9.5 pp, low confidence
GLM-4.769.4%6.2%estimated ± 9.5 pp, low confidence
GPT-5.2-Codex72.4%7.1%estimated ± 9.5 pp, low confidence
GPT-5.4 mini73.0%7.3%estimated ± 9.5 pp, low confidence
GPT-5.4 nano69.8%6.3%estimated ± 9.5 pp, low confidence
Grok 4.371.4%6.8%estimated ± 9.5 pp, low confidence
Kimi K2.7 Code78.2%9.4%estimated ± 9.5 pp, medium confidence
Mistral Medium 3.5 128B66.4%5.5%estimated ± 9.5 pp, low confidence
Nemotron 3 Ultra69.0%6.1%estimated ± 9.5 pp, low confidence
Qwen3.5 Flash64.4%5.0%estimated ± 9.5 pp, low confidence