benchgap
Calibration

SWE-bench Verified → FrontierCode 1.1 Main

FrontierCode 1.1 Main is estimated from SWE-bench Verified with a offset logistic curve fitted on 6 models measured on both: y = 0.0000 + (0.5456 − 0.0000) / (1 + exp(−24.83·(x − 0.8064))), R² = 0.99, cross-validated error 2.7 pp. It is used for 71 estimates.

Estimated modelSWE-bench VerifiedFrontierCode 1.1 MainSource
Apodex 1.177.7%17.7%estimated ± 2.7 pp, low confidence
BTL-478.4%19.9%estimated ± 2.7 pp, low confidence
Claude 3.5 Sonnet49.0%0.0%estimated ± 2.7 pp, low confidence
Claude 4.1 Opus74.5%9.7%estimated ± 2.7 pp, low confidence
Claude 4 Sonnet72.7%6.7%estimated ± 2.7 pp, low confidence
Claude Haiku 4.573.3%7.6%estimated ± 2.7 pp, low confidence
Claude Mythos 595.5%53.2%estimated ± 2.7 pp, medium confidence
Claude Opus 4.580.9%28.1%estimated ± 2.7 pp, medium confidence
Claude Opus 4.7 (Adaptive)87.6%46.3%estimated ± 2.7 pp, medium confidence
Claude Sonnet 4.577.2%16.3%estimated ± 2.7 pp, low confidence
DeepSeek V342.0%0.0%estimated ± 2.7 pp, low confidence
DeepSeek V4 Flash 073179.0%21.8%estimated ± 2.7 pp, low confidence
DeepSeek V4 Pro 081380.6%27.1%estimated ± 2.7 pp, medium confidence
dots3-note Preview78.4%19.9%estimated ± 2.7 pp, low confidence
Ember-192.2%51.6%estimated ± 2.7 pp, medium confidence
Gemini 2.5 Pro63.8%0.8%estimated ± 2.7 pp, low confidence
GLM-4.773.8%8.4%estimated ± 2.7 pp, low confidence
GLM-577.8%18.0%estimated ± 2.7 pp, low confidence
GPT-4.154.6%0.1%estimated ± 2.7 pp, low confidence
GPT-4.1 mini23.6%0.0%estimated ± 2.7 pp, low confidence
GPT-5.280.0%25.1%estimated ± 2.7 pp, medium confidence
GPT-5.3 Codex85.0%40.7%estimated ± 2.7 pp, medium confidence
Granite 4.2 30B57.0%0.2%estimated ± 2.7 pp, low confidence
Granite 4.2 8B47.7%0.0%estimated ± 2.7 pp, low confidence
Grok 4.2076.7%14.9%estimated ± 2.7 pp, low confidence
Grok Code Fast 170.8%4.4%estimated ± 2.7 pp, low confidence
Hy3 Preview74.4%9.5%estimated ± 2.7 pp, low confidence
Inkling77.6%17.4%estimated ± 2.7 pp, low confidence
Inkling-Small80.2%25.8%estimated ± 2.7 pp, medium confidence
K-EXAONE 2.068.2%2.4%estimated ± 2.7 pp, low confidence
Kimi K2.680.2%25.8%estimated ± 2.7 pp, medium confidence
Kimi K2.576.8%15.2%estimated ± 2.7 pp, low confidence
Kimi K2.5 (Reasoning)76.8%15.2%estimated ± 2.7 pp, low confidence
Laguna M.174.6%9.9%estimated ± 2.7 pp, low confidence
Laguna XS.269.9%3.5%estimated ± 2.7 pp, low confidence
Laguna XS 2.170.9%4.5%estimated ± 2.7 pp, low confidence
LLaDA2.2-flash49.3%0.0%estimated ± 2.7 pp, low confidence
LongCat-Flash-Lite-Sparse68.2%2.4%estimated ± 2.7 pp, low confidence
MAI-Code-1.1-Flash72.6%6.5%estimated ± 2.7 pp, low confidence
MAI-Thinking-173.5%7.9%estimated ± 2.7 pp, low confidence
MiMo-V2-Flash73.4%7.7%estimated ± 2.7 pp, low confidence
MiMo-V2-Omni74.8%10.4%estimated ± 2.7 pp, low confidence
MiMo-V2-Pro78.0%18.6%estimated ± 2.7 pp, low confidence
MiniCPM5-2B46.4%0.0%estimated ± 2.7 pp, low confidence
MiniMax M380.5%26.8%estimated ± 2.7 pp, medium confidence
Mistral Medium 3.5 128B77.6%17.4%estimated ± 2.7 pp, low confidence
Muse Glimmer 30B76.0%13.1%estimated ± 2.7 pp, low confidence
Muse Spark77.4%16.8%estimated ± 2.7 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP452.8%0.1%estimated ± 2.7 pp, low confidence
Nemotron 3 Ultra71.9%5.6%estimated ± 2.7 pp, low confidence
o3-mini49.3%0.0%estimated ± 2.7 pp, low confidence
Ornith-1.0-35B75.6%12.1%estimated ± 2.7 pp, low confidence
Ornith-1.0-397B82.4%33.1%estimated ± 2.7 pp, medium confidence
Ornith-1.0-9B69.4%3.2%estimated ± 2.7 pp, low confidence
Ornith-1.5-35B-A3B79.0%21.8%estimated ± 2.7 pp, low confidence
Ornith-1.5-397B86.0%43.1%estimated ± 2.7 pp, medium confidence
Ornith-1.5-9B70.6%4.2%estimated ± 2.7 pp, low confidence
Qwen3.5-122B-A10B72.0%5.7%estimated ± 2.7 pp, low confidence
Qwen3.5-27B72.4%6.2%estimated ± 2.7 pp, low confidence
Qwen3.5-35B-A3B69.2%3.0%estimated ± 2.7 pp, low confidence
Qwen3.5 397B76.2%13.6%estimated ± 2.7 pp, low confidence
Qwen3.6-27B77.2%16.3%estimated ± 2.7 pp, low confidence
Qwen3.6-35B-A3B73.4%7.7%estimated ± 2.7 pp, low confidence
Qwen3.6 Plus78.8%21.1%estimated ± 2.7 pp, low confidence
Qwen3.7 Max80.4%26.5%estimated ± 2.7 pp, medium confidence
Qwen3.7 Plus77.7%17.7%estimated ± 2.7 pp, low confidence
Beam80.9%28.1%estimated ± 2.7 pp, medium confidence
Solar Open 270.4%4.0%estimated ± 2.7 pp, low confidence
Solar Pro 470.6%4.2%estimated ± 2.7 pp, low confidence
Ternary Bonsai 2 27B60.8%0.4%estimated ± 2.7 pp, low confidence
ZAYA1-74B-Preview53.2%0.1%estimated ± 2.7 pp, low confidence