Calibration
SWE-bench Verified → LiveCodeBench
LiveCodeBench is estimated from SWE-bench Verified with a Michaelis–Menten curve fitted on 7 models measured on both: y = 2.0000·x / (1.04998 + x), R² = 0.77, cross-validated error 10.2 pp. It is used for 44 estimates.
| Estimated model | SWE-bench Verified | LiveCodeBench | Source |
|---|---|---|---|
| BTL-4 | 78.4% | 85.5% | estimated ± 10.2 pp, low confidence |
| Claude 3.5 Sonnet | 49.0% | 63.6% | estimated ± 10.2 pp, low confidence |
| Claude 4.1 Opus | 74.5% | 83.0% | estimated ± 10.2 pp, low confidence |
| Claude 4 Sonnet | 72.7% | 81.8% | estimated ± 10.2 pp, low confidence |
| Claude Haiku 4.5 | 73.3% | 82.2% | estimated ± 10.2 pp, low confidence |
| Claude Mythos 5 | 95.5% | 95.3% | estimated ± 10.2 pp, low confidence |
| Claude Opus 4.5 | 80.9% | 87.0% | estimated ± 10.2 pp, low confidence |
| Claude Opus 4.6 | 80.8% | 87.0% | estimated ± 10.2 pp, low confidence |
| Claude Sonnet 4.5 | 77.2% | 84.7% | estimated ± 10.2 pp, low confidence |
| Claude Sonnet 4.6 | 79.6% | 86.2% | estimated ± 10.2 pp, low confidence |
| dots3-note Preview | 78.4% | 85.5% | estimated ± 10.2 pp, low confidence |
| Ember-1 | 92.2% | 93.5% | estimated ± 10.2 pp, low confidence |
| GLM-5 | 77.8% | 85.1% | estimated ± 10.2 pp, low confidence |
| GPT-4.1 | 54.6% | 68.4% | estimated ± 10.2 pp, low confidence |
| GPT-5.2 | 80.0% | 86.5% | estimated ± 10.2 pp, low confidence |
| GPT-5.3 Codex | 85.0% | 89.5% | estimated ± 10.2 pp, low confidence |
| Granite 4.2 30B | 57.0% | 70.4% | estimated ± 10.2 pp, low confidence |
| Grok 4.20 | 76.7% | 84.4% | estimated ± 10.2 pp, low confidence |
| Grok Code Fast 1 | 70.8% | 80.5% | estimated ± 10.2 pp, low confidence |
| K-EXAONE 2.0 | 68.2% | 78.8% | estimated ± 10.2 pp, low confidence |
| Laguna M.1 | 74.6% | 83.1% | estimated ± 10.2 pp, low confidence |
| Laguna XS.2 | 69.9% | 79.9% | estimated ± 10.2 pp, low confidence |
| Laguna XS 2.1 | 70.9% | 80.6% | estimated ± 10.2 pp, low confidence |
| LLaDA2.2-flash | 49.3% | 63.9% | estimated ± 10.2 pp, low confidence |
| LongCat-Flash-Lite-Sparse | 68.2% | 78.8% | estimated ± 10.2 pp, low confidence |
| MAI-Code-1.1-Flash | 72.6% | 81.8% | estimated ± 10.2 pp, low confidence |
| MAI-Thinking-1 | 73.5% | 82.4% | estimated ± 10.2 pp, low confidence |
| MiMo-V2-Omni | 74.8% | 83.2% | estimated ± 10.2 pp, low confidence |
| MiMo-V2-Pro | 78.0% | 85.2% | estimated ± 10.2 pp, low confidence |
| MiniCPM5-2B | 46.4% | 61.3% | estimated ± 10.2 pp, low confidence |
| o3-mini | 49.3% | 63.9% | estimated ± 10.2 pp, low confidence |
| Ornith-1.0-35B | 75.6% | 83.7% | estimated ± 10.2 pp, low confidence |
| Ornith-1.0-397B | 82.4% | 87.9% | estimated ± 10.2 pp, low confidence |
| Ornith-1.0-9B | 69.4% | 79.6% | estimated ± 10.2 pp, low confidence |
| Ornith-1.5-35B-A3B | 79.0% | 85.9% | estimated ± 10.2 pp, low confidence |
| Ornith-1.5-397B | 86.0% | 90.1% | estimated ± 10.2 pp, low confidence |
| Ornith-1.5-9B | 70.6% | 80.4% | estimated ± 10.2 pp, low confidence |
| Qwen3.5-27B | 72.4% | 81.6% | estimated ± 10.2 pp, low confidence |
| Qwen3.5-35B-A3B | 69.2% | 79.4% | estimated ± 10.2 pp, low confidence |
| Qwen3.5 397B | 76.2% | 84.1% | estimated ± 10.2 pp, low confidence |
| Beam | 80.9% | 87.0% | estimated ± 10.2 pp, low confidence |
| Solar Open 2 | 70.4% | 80.3% | estimated ± 10.2 pp, low confidence |
| Ternary Bonsai 2 27B | 60.8% | 73.3% | estimated ± 10.2 pp, low confidence |
| ZAYA1-74B-Preview | 53.2% | 67.3% | estimated ± 10.2 pp, low confidence |