Calibration
DeepSWE → SWE-bench Verified
SWE-bench Verified is estimated from DeepSWE with a Hill curve fitted on 7 models measured on both: y = 0.7866 + (1.1365 − 0.7866)·x^6.00 / (0.78061^6.00 + x^6.00), R² = 0.67, cross-validated error 5.2 pp. It is used for 10 estimates.
| Estimated model | DeepSWE | SWE-bench Verified | Source |
|---|---|---|---|
| GPT-6.1 Sol | 71.9% | 91.9% | estimated ± 5.2 pp, low confidence |
| GPT-6 Luna | 66.6% | 88.4% | estimated ± 5.2 pp, low confidence |
| GPT-6 Sol | 68.8% | 89.8% | estimated ± 5.2 pp, low confidence |
| Grok 4.7 | 71.0% | 91.3% | estimated ± 5.2 pp, low confidence |
| MiMo-V2.6-Flash | 67.9% | 89.2% | estimated ± 5.2 pp, low confidence |
| MiMo-V2.6-Pro | 71.9% | 91.9% | estimated ± 5.2 pp, low confidence |
| Muse Spark 1.3 | 75.4% | 94.3% | estimated ± 5.2 pp, low confidence |
| Pareto 26.10 Preview | 69.9% | 90.6% | estimated ± 5.2 pp, low confidence |
| Pareto 26.9 | 74.0% | 93.4% | estimated ± 5.2 pp, low confidence |
| Step 5 Preview | 67.7% | 89.1% | estimated ± 5.2 pp, low confidence |