benchgap
Calibration

SWE-bench Pro → FrontierSWE v2

FrontierSWE v2 is estimated from SWE-bench Pro with a linear curve fitted on 8 models measured on both: y = 1.7808·x + -0.9171, R² = 0.88, cross-validated error 9.0 pp. It is used for 62 estimates.

Estimated modelSWE-bench ProFrontierSWE v2Source
Atria Dawn Preview59.6%14.4%estimated ± 9.0 pp, medium confidence
Claude Mythos 580.3%51.3%estimated ± 9.0 pp, medium confidence
Claude Opus 4.557.1%10.0%estimated ± 9.0 pp, medium confidence
Claude Opus 4.653.4%3.4%estimated ± 9.0 pp, low confidence
Claude Opus 4.7 (Adaptive)64.3%22.8%estimated ± 9.0 pp, medium confidence
Claude Opus 4.869.2%31.5%estimated ± 9.0 pp, medium confidence
DeepSeek V4 Flash 073152.6%2.0%estimated ± 9.0 pp, low confidence
DeepSeek V4 Pro 081355.4%6.9%estimated ± 9.0 pp, medium confidence
dots3-note Preview61.0%16.9%estimated ± 9.0 pp, medium confidence
Gemini 3.5 Flash55.1%6.4%estimated ± 9.0 pp, medium confidence
Gemini 3.5 Flash-Lite54.2%4.8%estimated ± 9.0 pp, low confidence
GLM-555.1%6.4%estimated ± 9.0 pp, medium confidence
GLM-5.158.4%12.3%estimated ± 9.0 pp, medium confidence
GLM-5.262.1%18.9%estimated ± 9.0 pp, medium confidence
GPT-5.255.6%7.3%estimated ± 9.0 pp, medium confidence
GPT-5.3 Codex56.8%9.4%estimated ± 9.0 pp, medium confidence
GPT-5.457.7%11.0%estimated ± 9.0 pp, medium confidence
GPT-5.558.6%12.6%estimated ± 9.0 pp, medium confidence
Granite 4.2 30B33.3%0.0%estimated ± 9.0 pp, low confidence
Granite 4.2 8B19.1%0.0%estimated ± 9.0 pp, low confidence
Grok 4.2051.8%0.5%estimated ± 9.0 pp, low confidence
Grok 4.564.7%23.5%estimated ± 9.0 pp, medium confidence
Hy4 preview65.7%25.3%estimated ± 9.0 pp, medium confidence
Inkling-Small55.9%7.8%estimated ± 9.0 pp, medium confidence
Kimi K2.658.6%12.6%estimated ± 9.0 pp, medium confidence
Kimi K2.550.7%0.0%estimated ± 9.0 pp, low confidence
Laguna M.149.2%0.0%estimated ± 9.0 pp, low confidence
Laguna S 2.159.4%14.1%estimated ± 9.0 pp, medium confidence
Laguna XS.246.3%0.0%estimated ± 9.0 pp, low confidence
Laguna XS 2.147.6%0.0%estimated ± 9.0 pp, low confidence
Ling 3.0 Flash56.6%9.1%estimated ± 9.0 pp, medium confidence
LLaDA2.2-flash30.1%0.0%estimated ± 9.0 pp, low confidence
LongCat-Flash-Lite-Sparse40.6%0.0%estimated ± 9.0 pp, low confidence
MAI-Thinking-152.8%2.3%estimated ± 9.0 pp, low confidence
MiMo-V2.556.1%8.2%estimated ± 9.0 pp, medium confidence
MiMo-V2.5-Pro57.2%10.2%estimated ± 9.0 pp, medium confidence
MiniCPM5-2B14.4%0.0%estimated ± 9.0 pp, low confidence
MiniMax M2.756.2%8.4%estimated ± 9.0 pp, medium confidence
MiniMax M359.0%13.4%estimated ± 9.0 pp, medium confidence
Muse Glimmer 30B51.2%0.0%estimated ± 9.0 pp, low confidence
Muse Spark52.4%1.6%estimated ± 9.0 pp, low confidence
Muse Spark 1.161.5%17.8%estimated ± 9.0 pp, medium confidence
Ornith-1.0-35B50.4%0.0%estimated ± 9.0 pp, low confidence
Ornith-1.0-397B62.2%19.1%estimated ± 9.0 pp, medium confidence
Ornith-1.0-9B42.9%0.0%estimated ± 9.0 pp, low confidence
Ornith-1.5-35B-A3B59.6%14.4%estimated ± 9.0 pp, medium confidence
Ornith-1.5-397B65.1%24.2%estimated ± 9.0 pp, medium confidence
Ornith-1.5-9B47.5%0.0%estimated ± 9.0 pp, low confidence
Qwen3.5 397B50.9%0.0%estimated ± 9.0 pp, low confidence
Qwen3.6-27B53.5%3.6%estimated ± 9.0 pp, low confidence
Qwen3.6-35B-A3B49.5%0.0%estimated ± 9.0 pp, low confidence
Qwen 3.6 Max (preview)57.3%10.3%estimated ± 9.0 pp, medium confidence
Qwen3.6 Plus56.6%9.1%estimated ± 9.0 pp, medium confidence
Qwen3.7 Max60.6%16.2%estimated ± 9.0 pp, medium confidence
Qwen3.7 Plus57.6%10.9%estimated ± 9.0 pp, medium confidence
Qwen3.8-27B61.7%18.2%estimated ± 9.0 pp, medium confidence
Qwen3.8-Flash-Next62.5%19.6%estimated ± 9.0 pp, medium confidence
Qwen3.8-Omni-Flash63.3%21.0%estimated ± 9.0 pp, medium confidence
Beam65.5%24.9%estimated ± 9.0 pp, medium confidence
Sakana Fugu59.0%13.4%estimated ± 9.0 pp, medium confidence
Sakana Fugu-Ultra73.7%39.5%estimated ± 9.0 pp, medium confidence
Step 3.7 Flash56.3%8.6%estimated ± 9.0 pp, medium confidence