benchgap
Calibration

Vals SWE-bench → FrontierCode 1.1 Extended

FrontierCode 1.1 Extended is estimated from Vals SWE-bench with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.95815^6.00 + x^6.00), R² = 0.69, cross-validated error 2.4 pp. It is used for 43 estimates.

Estimated modelVals SWE-benchFrontierCode 1.1 ExtendedSource
Claude Haiku 4.566.6%12.2%estimated ± 2.4 pp, low confidence
Claude Opus 4.782.0%33.8%estimated ± 2.4 pp, low confidence
Claude Sonnet 4.677.4%26.1%estimated ± 2.4 pp, low confidence
DeepSeek V4 Flash 073188.8%46.5%estimated ± 2.4 pp, low confidence
DeepSeek V4 Pro 081396.4%61.1%estimated ± 2.4 pp, medium confidence
Gemini 2.5 Pro54.4%3.9%estimated ± 2.4 pp, low confidence
Gemini 3.1 Flash-Lite62.8%8.8%estimated ± 2.4 pp, low confidence
Gemini 3.1 Pro78.8%28.4%estimated ± 2.4 pp, low confidence
Gemini 3.5 Flash-Lite75.0%22.4%estimated ± 2.4 pp, low confidence
Gemini 3.7 Flash80.8%31.7%estimated ± 2.4 pp, low confidence
Gemini 3 Flash75.0%22.4%estimated ± 2.4 pp, low confidence
GLM-4.769.4%15.1%estimated ± 2.4 pp, low confidence
GLM-5.176.4%24.5%estimated ± 2.4 pp, low confidence
GLM-5.395.4%59.2%estimated ± 2.4 pp, medium confidence
GLM-5.3-Flash92.0%52.7%estimated ± 2.4 pp, low confidence
GPT-5.2-Codex72.4%18.8%estimated ± 2.4 pp, low confidence
GPT-5.3 Codex78.0%27.1%estimated ± 2.4 pp, low confidence
GPT-5.4 mini73.0%19.6%estimated ± 2.4 pp, low confidence
GPT-5.4 nano69.8%15.6%estimated ± 2.4 pp, low confidence
Grok 4.2072.2%18.6%estimated ± 2.4 pp, low confidence
Grok 4.371.4%17.5%estimated ± 2.4 pp, low confidence
Inkling77.6%26.4%estimated ± 2.4 pp, low confidence
Inkling-Small82.2%34.2%estimated ± 2.4 pp, low confidence
Kimi K2.676.2%24.2%estimated ± 2.4 pp, low confidence
Laguna M.157.6%5.4%estimated ± 2.4 pp, low confidence
Laguna XS.255.2%4.2%estimated ± 2.4 pp, low confidence
Ling 3.0 Flash65.2%10.8%estimated ± 2.4 pp, low confidence
MiMo-V2.571.0%17.0%estimated ± 2.4 pp, low confidence
MiMo-V2.5-Pro74.0%21.0%estimated ± 2.4 pp, low confidence
MiniMax M2.773.8%20.7%estimated ± 2.4 pp, low confidence
MiniMax M375.0%22.4%estimated ± 2.4 pp, low confidence
Mistral Medium 3.5 128B66.4%12.0%estimated ± 2.4 pp, low confidence
Muse Spark74.4%21.6%estimated ± 2.4 pp, low confidence
Muse Spark 1.182.0%33.8%estimated ± 2.4 pp, low confidence
Muse Spark 1.286.6%42.3%estimated ± 2.4 pp, low confidence
Nemotron 3 Ultra69.0%14.7%estimated ± 2.4 pp, low confidence
Qwen3.5 Flash64.4%10.1%estimated ± 2.4 pp, low confidence
Qwen3.6-27B70.0%15.8%estimated ± 2.4 pp, low confidence
Qwen 3.6 Max (preview)72.8%19.4%estimated ± 2.4 pp, low confidence
Qwen3.6 Plus73.4%20.2%estimated ± 2.4 pp, low confidence
Qwen3.7 Max68.8%14.5%estimated ± 2.4 pp, low confidence
Qwen3.8-27B86.0%41.2%estimated ± 2.4 pp, low confidence
Qwen3.8 Max85.6%40.4%estimated ± 2.4 pp, low confidence