Calibration
DeepSWE → PostTrainBench v1.1
PostTrainBench v1.1 is estimated from DeepSWE with a linear curve fitted on 8 models measured on both: y = 0.9562·x + -0.2819, R² = 0.76, cross-validated error 4.8 pp. It is used for 12 estimates.
| Estimated model | DeepSWE | PostTrainBench v1.1 | Source |
|---|---|---|---|
| DeepSeek V4.1 Flash | 74.2% | 42.8% | estimated ± 4.8 pp, high confidence |
| Ember-1 | 75.2% | 43.7% | estimated ± 4.8 pp, high confidence |
| GPT-6.1 Sol | 71.9% | 40.6% | estimated ± 4.8 pp, high confidence |
| GPT-6 Luna | 66.6% | 35.5% | estimated ± 4.8 pp, high confidence |
| GPT-6 Sol | 68.8% | 37.6% | estimated ± 4.8 pp, high confidence |
| MiMo-V2.6-Flash | 67.9% | 36.7% | estimated ± 4.8 pp, high confidence |
| MiMo-V2.6-Pro | 71.9% | 40.6% | estimated ± 4.8 pp, high confidence |
| Muse Spark 1.3 | 75.4% | 43.9% | estimated ± 4.8 pp, high confidence |
| Pareto 26.10 Preview | 69.9% | 38.7% | estimated ± 4.8 pp, high confidence |
| Pareto 26.9 | 74.0% | 42.6% | estimated ± 4.8 pp, high confidence |
| Step 5 Preview | 67.7% | 36.5% | estimated ± 4.8 pp, high confidence |
| SWE-2 | 73.0% | 41.6% | estimated ± 4.8 pp, high confidence |