Calibration
SWE-bench Verified → FrontierCode 1.1 Main
FrontierCode 1.1 Main is estimated from SWE-bench Verified with a offset logistic curve fitted on 6 models measured on both: y = 0.0000 + (0.5456 − 0.0000) / (1 + exp(−24.83·(x − 0.8064))), R² = 0.99, cross-validated error 2.7 pp. It is used for 71 estimates.
| Estimated model | SWE-bench Verified | FrontierCode 1.1 Main | Source |
|---|---|---|---|
| Apodex 1.1 | 77.7% | 17.7% | estimated ± 2.7 pp, low confidence |
| BTL-4 | 78.4% | 19.9% | estimated ± 2.7 pp, low confidence |
| Claude 3.5 Sonnet | 49.0% | 0.0% | estimated ± 2.7 pp, low confidence |
| Claude 4.1 Opus | 74.5% | 9.7% | estimated ± 2.7 pp, low confidence |
| Claude 4 Sonnet | 72.7% | 6.7% | estimated ± 2.7 pp, low confidence |
| Claude Haiku 4.5 | 73.3% | 7.6% | estimated ± 2.7 pp, low confidence |
| Claude Mythos 5 | 95.5% | 53.2% | estimated ± 2.7 pp, medium confidence |
| Claude Opus 4.5 | 80.9% | 28.1% | estimated ± 2.7 pp, medium confidence |
| Claude Opus 4.7 (Adaptive) | 87.6% | 46.3% | estimated ± 2.7 pp, medium confidence |
| Claude Sonnet 4.5 | 77.2% | 16.3% | estimated ± 2.7 pp, low confidence |
| DeepSeek V3 | 42.0% | 0.0% | estimated ± 2.7 pp, low confidence |
| DeepSeek V4 Flash 0731 | 79.0% | 21.8% | estimated ± 2.7 pp, low confidence |
| DeepSeek V4 Pro 0813 | 80.6% | 27.1% | estimated ± 2.7 pp, medium confidence |
| dots3-note Preview | 78.4% | 19.9% | estimated ± 2.7 pp, low confidence |
| Ember-1 | 92.2% | 51.6% | estimated ± 2.7 pp, medium confidence |
| Gemini 2.5 Pro | 63.8% | 0.8% | estimated ± 2.7 pp, low confidence |
| GLM-4.7 | 73.8% | 8.4% | estimated ± 2.7 pp, low confidence |
| GLM-5 | 77.8% | 18.0% | estimated ± 2.7 pp, low confidence |
| GPT-4.1 | 54.6% | 0.1% | estimated ± 2.7 pp, low confidence |
| GPT-4.1 mini | 23.6% | 0.0% | estimated ± 2.7 pp, low confidence |
| GPT-5.2 | 80.0% | 25.1% | estimated ± 2.7 pp, medium confidence |
| GPT-5.3 Codex | 85.0% | 40.7% | estimated ± 2.7 pp, medium confidence |
| Granite 4.2 30B | 57.0% | 0.2% | estimated ± 2.7 pp, low confidence |
| Granite 4.2 8B | 47.7% | 0.0% | estimated ± 2.7 pp, low confidence |
| Grok 4.20 | 76.7% | 14.9% | estimated ± 2.7 pp, low confidence |
| Grok Code Fast 1 | 70.8% | 4.4% | estimated ± 2.7 pp, low confidence |
| Hy3 Preview | 74.4% | 9.5% | estimated ± 2.7 pp, low confidence |
| Inkling | 77.6% | 17.4% | estimated ± 2.7 pp, low confidence |
| Inkling-Small | 80.2% | 25.8% | estimated ± 2.7 pp, medium confidence |
| K-EXAONE 2.0 | 68.2% | 2.4% | estimated ± 2.7 pp, low confidence |
| Kimi K2.6 | 80.2% | 25.8% | estimated ± 2.7 pp, medium confidence |
| Kimi K2.5 | 76.8% | 15.2% | estimated ± 2.7 pp, low confidence |
| Kimi K2.5 (Reasoning) | 76.8% | 15.2% | estimated ± 2.7 pp, low confidence |
| Laguna M.1 | 74.6% | 9.9% | estimated ± 2.7 pp, low confidence |
| Laguna XS.2 | 69.9% | 3.5% | estimated ± 2.7 pp, low confidence |
| Laguna XS 2.1 | 70.9% | 4.5% | estimated ± 2.7 pp, low confidence |
| LLaDA2.2-flash | 49.3% | 0.0% | estimated ± 2.7 pp, low confidence |
| LongCat-Flash-Lite-Sparse | 68.2% | 2.4% | estimated ± 2.7 pp, low confidence |
| MAI-Code-1.1-Flash | 72.6% | 6.5% | estimated ± 2.7 pp, low confidence |
| MAI-Thinking-1 | 73.5% | 7.9% | estimated ± 2.7 pp, low confidence |
| MiMo-V2-Flash | 73.4% | 7.7% | estimated ± 2.7 pp, low confidence |
| MiMo-V2-Omni | 74.8% | 10.4% | estimated ± 2.7 pp, low confidence |
| MiMo-V2-Pro | 78.0% | 18.6% | estimated ± 2.7 pp, low confidence |
| MiniCPM5-2B | 46.4% | 0.0% | estimated ± 2.7 pp, low confidence |
| MiniMax M3 | 80.5% | 26.8% | estimated ± 2.7 pp, medium confidence |
| Mistral Medium 3.5 128B | 77.6% | 17.4% | estimated ± 2.7 pp, low confidence |
| Muse Glimmer 30B | 76.0% | 13.1% | estimated ± 2.7 pp, low confidence |
| Muse Spark | 77.4% | 16.8% | estimated ± 2.7 pp, low confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 52.8% | 0.1% | estimated ± 2.7 pp, low confidence |
| Nemotron 3 Ultra | 71.9% | 5.6% | estimated ± 2.7 pp, low confidence |
| o3-mini | 49.3% | 0.0% | estimated ± 2.7 pp, low confidence |
| Ornith-1.0-35B | 75.6% | 12.1% | estimated ± 2.7 pp, low confidence |
| Ornith-1.0-397B | 82.4% | 33.1% | estimated ± 2.7 pp, medium confidence |
| Ornith-1.0-9B | 69.4% | 3.2% | estimated ± 2.7 pp, low confidence |
| Ornith-1.5-35B-A3B | 79.0% | 21.8% | estimated ± 2.7 pp, low confidence |
| Ornith-1.5-397B | 86.0% | 43.1% | estimated ± 2.7 pp, medium confidence |
| Ornith-1.5-9B | 70.6% | 4.2% | estimated ± 2.7 pp, low confidence |
| Qwen3.5-122B-A10B | 72.0% | 5.7% | estimated ± 2.7 pp, low confidence |
| Qwen3.5-27B | 72.4% | 6.2% | estimated ± 2.7 pp, low confidence |
| Qwen3.5-35B-A3B | 69.2% | 3.0% | estimated ± 2.7 pp, low confidence |
| Qwen3.5 397B | 76.2% | 13.6% | estimated ± 2.7 pp, low confidence |
| Qwen3.6-27B | 77.2% | 16.3% | estimated ± 2.7 pp, low confidence |
| Qwen3.6-35B-A3B | 73.4% | 7.7% | estimated ± 2.7 pp, low confidence |
| Qwen3.6 Plus | 78.8% | 21.1% | estimated ± 2.7 pp, low confidence |
| Qwen3.7 Max | 80.4% | 26.5% | estimated ± 2.7 pp, medium confidence |
| Qwen3.7 Plus | 77.7% | 17.7% | estimated ± 2.7 pp, low confidence |
| Beam | 80.9% | 28.1% | estimated ± 2.7 pp, medium confidence |
| Solar Open 2 | 70.4% | 4.0% | estimated ± 2.7 pp, low confidence |
| Solar Pro 4 | 70.6% | 4.2% | estimated ± 2.7 pp, low confidence |
| Ternary Bonsai 2 27B | 60.8% | 0.4% | estimated ± 2.7 pp, low confidence |
| ZAYA1-74B-Preview | 53.2% | 0.1% | estimated ± 2.7 pp, low confidence |