Calibration
SWE-bench Verified → LiveCodeBench v6
LiveCodeBench v6 is estimated from SWE-bench Verified with a Hill curve fitted on 15 models measured on both: y = 0.5362 + (0.8783 − 0.5362)·x^6.00 / (0.48357^6.00 + x^6.00), R² = 0.42, cross-validated error 8.1 pp. It is used for 15 estimates.
| Estimated model | SWE-bench Verified | LiveCodeBench v6 | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 49.0% | 71.4% | estimated ± 8.1 pp, low confidence |
| Claude 4.1 Opus | 74.5% | 85.4% | estimated ± 8.1 pp, low confidence |
| Claude 4 Sonnet | 72.7% | 85.1% | estimated ± 8.1 pp, low confidence |
| Claude Haiku 4.5 | 73.3% | 85.2% | estimated ± 8.1 pp, low confidence |
| Claude Sonnet 4.5 | 77.2% | 85.9% | estimated ± 8.1 pp, low confidence |
| Claude Sonnet 4.6 | 79.6% | 86.2% | estimated ± 8.1 pp, low confidence |
| Ember-1 | 92.2% | 87.1% | estimated ± 8.1 pp, low confidence |
| GPT-4.1 | 54.6% | 76.7% | estimated ± 8.1 pp, low confidence |
| Grok Code Fast 1 | 70.8% | 84.7% | estimated ± 8.1 pp, low confidence |
| MAI-Code-1.1-Flash | 72.6% | 85.1% | estimated ± 8.1 pp, low confidence |
| MiMo-V2-Omni | 74.8% | 85.5% | estimated ± 8.1 pp, low confidence |
| MiMo-V2-Pro | 78.0% | 86.0% | estimated ± 8.1 pp, low confidence |
| o3-mini | 49.3% | 71.7% | estimated ± 8.1 pp, low confidence |
| Qwen3.5-27B | 72.4% | 85.0% | estimated ± 8.1 pp, low confidence |
| Qwen3.5-35B-A3B | 69.2% | 84.3% | estimated ± 8.1 pp, low confidence |