Calibration
SWE-bench Verified → NL2Repo
NL2Repo is estimated from SWE-bench Verified with a Hill curve fitted on 15 models measured on both: y = 0.0131 + (1.1525 − 0.0131)·x^6.00 / (0.85375^6.00 + x^6.00), R² = 0.70, cross-validated error 6.1 pp. It is used for 23 estimates.
| Estimated model | SWE-bench Verified | NL2Repo | Source |
|---|---|---|---|
| Claude 3.5 Sonnet | 49.0% | 5.2% | estimated ± 6.1 pp, low confidence |
| Claude 4.1 Opus | 74.5% | 36.2% | estimated ± 6.1 pp, medium confidence |
| Claude 4 Sonnet | 72.7% | 32.8% | estimated ± 6.1 pp, medium confidence |
| Claude Mythos 5 | 95.5% | 76.7% | estimated ± 6.1 pp, low confidence |
| Claude Opus 4.6 | 80.8% | 49.0% | estimated ± 6.1 pp, medium confidence |
| Claude Sonnet 4.5 | 77.2% | 41.6% | estimated ± 6.1 pp, medium confidence |
| Ember-1 | 92.2% | 71.2% | estimated ± 6.1 pp, low confidence |
| GLM-5 | 77.8% | 42.8% | estimated ± 6.1 pp, medium confidence |
| GPT-4.1 | 54.6% | 8.6% | estimated ± 6.1 pp, low confidence |
| GPT-5.2 | 80.0% | 47.3% | estimated ± 6.1 pp, medium confidence |
| Grok Code Fast 1 | 70.8% | 29.3% | estimated ± 6.1 pp, medium confidence |
| K-EXAONE 2.0 | 68.2% | 24.8% | estimated ± 6.1 pp, low confidence |
| Laguna XS 2.1 | 70.9% | 29.5% | estimated ± 6.1 pp, medium confidence |
| LLaDA2.2-flash | 49.3% | 5.4% | estimated ± 6.1 pp, low confidence |
| LongCat-Flash-Lite-Sparse | 68.2% | 24.8% | estimated ± 6.1 pp, low confidence |
| MAI-Code-1.1-Flash | 72.6% | 32.6% | estimated ± 6.1 pp, medium confidence |
| MiMo-V2-Omni | 74.8% | 36.8% | estimated ± 6.1 pp, medium confidence |
| MiMo-V2-Pro | 78.0% | 43.2% | estimated ± 6.1 pp, medium confidence |
| o3-mini | 49.3% | 5.4% | estimated ± 6.1 pp, low confidence |
| Qwen3.5-27B | 72.4% | 32.2% | estimated ± 6.1 pp, medium confidence |
| Qwen3.5-35B-A3B | 69.2% | 26.5% | estimated ± 6.1 pp, low confidence |
| Beam | 80.9% | 49.2% | estimated ± 6.1 pp, medium confidence |
| Solar Pro 4 | 70.6% | 28.9% | estimated ± 6.1 pp, medium confidence |