Calibration
SWE-bench Verified → SWE-bench Pro
SWE-bench Pro is estimated from SWE-bench Verified with a inverse Michaelis–Menten curve fitted on 45 models measured on both: y = 1.14763·(x − 0.1358) / (0.1358 + 2.0000 − x), R² = 0.95, cross-validated error 3.4 pp. It is used for 43 estimates.
| Estimated model | SWE-bench Verified | SWE-bench Pro | Source |
|---|---|---|---|
| Ember-1 | 92.2% | 74.3% | estimated ± 3.4 pp, high confidence |
| Claude Sonnet 4.6 | 79.6% | 56.5% | estimated ± 3.4 pp, high confidence |
| BTL-4 | 78.4% | 55.0% | estimated ± 3.4 pp, high confidence |
| MiMo-V2-Pro | 78.0% | 54.5% | estimated ± 3.4 pp, high confidence |
| Apodex 1.1 | 77.7% | 54.2% | estimated ± 3.4 pp, high confidence |
| Mistral Medium 3.5 128B | 77.6% | 54.0% | estimated ± 3.4 pp, high confidence |
| Claude Sonnet 4.5 | 77.2% | 53.5% | estimated ± 3.4 pp, high confidence |
| MiMo-V2-Omni | 74.8% | 50.6% | estimated ± 3.4 pp, high confidence |
| Claude 4.1 Opus | 74.5% | 50.3% | estimated ± 3.4 pp, high confidence |
| Hy3 Preview | 74.4% | 50.1% | estimated ± 3.4 pp, high confidence |
| Step 3.5 Flash | 74.4% | 50.1% | estimated ± 3.4 pp, high confidence |
| GLM-4.7 | 73.8% | 49.4% | estimated ± 3.4 pp, high confidence |
| MiMo-V2-Flash | 73.4% | 49.0% | estimated ± 3.4 pp, high confidence |
| Claude Haiku 4.5 | 73.3% | 48.9% | estimated ± 3.4 pp, high confidence |
| Claude 4 Sonnet | 72.7% | 48.2% | estimated ± 3.4 pp, high confidence |
| MAI-Code-1.1-Flash | 72.6% | 48.0% | estimated ± 3.4 pp, high confidence |
| Qwen3.5-27B | 72.4% | 47.8% | estimated ± 3.4 pp, high confidence |
| Qwen3.5-122B-A10B | 72.0% | 47.3% | estimated ± 3.4 pp, high confidence |
| Nemotron 3 Ultra | 71.9% | 47.2% | estimated ± 3.4 pp, high confidence |
| Grok Code Fast 1 | 70.8% | 46.0% | estimated ± 3.4 pp, high confidence |
| Solar Pro 4 | 70.6% | 45.8% | estimated ± 3.4 pp, high confidence |
| Solar Open 2 | 70.4% | 45.5% | estimated ± 3.4 pp, high confidence |
| Qwen3.5-35B-A3B | 69.2% | 44.2% | estimated ± 3.4 pp, high confidence |
| K-EXAONE 2.0 | 68.2% | 43.1% | estimated ± 3.4 pp, high confidence |
| DeepSeek V3.1 | 66.0% | 40.8% | estimated ± 3.4 pp, high confidence |
| Kimi K2 | 65.8% | 40.5% | estimated ± 3.4 pp, high confidence |
| GPT-OSS 120B | 62.4% | 37.1% | estimated ± 3.4 pp, high confidence |
| Ternary Bonsai 2 27B | 60.8% | 35.5% | estimated ± 3.4 pp, high confidence |
| GPT-OSS 20B | 60.7% | 35.4% | estimated ± 3.4 pp, high confidence |
| Nemotron 3 Super 100B | 60.5% | 35.1% | estimated ± 3.4 pp, high confidence |
| GLM-4.7-Flash | 59.2% | 33.9% | estimated ± 3.4 pp, high confidence |
| MiniMax M1 80k | 56.0% | 30.9% | estimated ± 3.4 pp, high confidence |
| GPT-4.1 | 54.6% | 29.6% | estimated ± 3.4 pp, high confidence |
| ZAYA1-74B-Preview | 53.2% | 28.3% | estimated ± 3.4 pp, high confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 52.8% | 28.0% | estimated ± 3.4 pp, high confidence |
| K-Exaone | 49.4% | 25.0% | estimated ± 3.4 pp, high confidence |
| o3-mini | 49.3% | 25.0% | estimated ± 3.4 pp, high confidence |
| DeepSeek-R1 | 49.2% | 24.9% | estimated ± 3.4 pp, high confidence |
| Claude 3.5 Sonnet | 49.0% | 24.7% | estimated ± 3.4 pp, high confidence |
| Sarvam 105B | 45.0% | 21.4% | estimated ± 3.4 pp, medium confidence |
| DeepSeek V3 | 42.0% | 19.0% | estimated ± 3.4 pp, medium confidence |
| Sarvam 30B | 34.0% | 13.0% | estimated ± 3.4 pp, medium confidence |
| GPT-4.1 mini | 23.6% | 6.1% | estimated ± 3.4 pp, medium confidence |