Calibration
SWE-bench Verified → Vals LiveCodeBench
Vals LiveCodeBench is estimated from SWE-bench Verified with a Hill curve fitted on 21 models measured on both: y = 0.0000 + (1.0017 − 0.0000)·x^6.00 / (0.61912^6.00 + x^6.00), R² = 0.39, cross-validated error 9.9 pp. It is used for 23 estimates.
| Estimated model | SWE-bench Verified | Vals LiveCodeBench | Source |
|---|---|---|---|
| Ember-1 | 92.2% | 91.8% | estimated ± 9.9 pp, low confidence |
| MiMo-V2-Pro | 78.0% | 80.1% | estimated ± 9.9 pp, low confidence |
| Apodex 1.1 | 77.7% | 79.8% | estimated ± 9.9 pp, low confidence |
| Mistral Medium 3.5 128B | 77.6% | 79.6% | estimated ± 9.9 pp, low confidence |
| Claude Sonnet 4.5 | 77.2% | 79.1% | estimated ± 9.9 pp, low confidence |
| MiMo-V2-Omni | 74.8% | 75.8% | estimated ± 9.9 pp, low confidence |
| Claude 4.1 Opus | 74.5% | 75.3% | estimated ± 9.9 pp, low confidence |
| Claude 4 Sonnet | 72.7% | 72.5% | estimated ± 9.9 pp, low confidence |
| MAI-Code-1.1-Flash | 72.6% | 72.3% | estimated ± 9.9 pp, low confidence |
| Solar Pro 4 | 70.6% | 68.9% | estimated ± 9.9 pp, low confidence |
| DeepSeek V3.1 | 66.0% | 59.6% | estimated ± 9.9 pp, low confidence |
| Gemini 2.5 Pro | 63.8% | 54.6% | estimated ± 9.9 pp, low confidence |
| Nemotron 3 Super 100B | 60.5% | 46.5% | estimated ± 9.9 pp, low confidence |
| MiniMax M1 80k | 56.0% | 35.4% | estimated ± 9.9 pp, low confidence |
| GPT-4.1 | 54.6% | 32.0% | estimated ± 9.9 pp, low confidence |
| K-Exaone | 49.4% | 20.5% | estimated ± 9.9 pp, low confidence |
| o3-mini | 49.3% | 20.3% | estimated ± 9.9 pp, low confidence |
| DeepSeek-R1 | 49.2% | 20.2% | estimated ± 9.9 pp, low confidence |
| Claude 3.5 Sonnet | 49.0% | 19.8% | estimated ± 9.9 pp, low confidence |
| Sarvam 105B | 45.0% | 12.9% | estimated ± 9.9 pp, low confidence |
| DeepSeek V3 | 42.0% | 8.9% | estimated ± 9.9 pp, low confidence |
| Sarvam 30B | 34.0% | 2.7% | estimated ± 9.9 pp, low confidence |
| GPT-4.1 mini | 23.6% | 0.3% | estimated ± 9.9 pp, low confidence |