Calibration
AA-SciCode → Vals SWE-bench
Vals SWE-bench is estimated from AA-SciCode with a linear curve fitted on 42 models measured on both: y = 1.2595·x + 0.1627, R² = 0.47, cross-validated error 7.6 pp. It is used for 11 estimates.
| Estimated model | AA-SciCode | Vals SWE-bench | Source |
|---|---|---|---|
| DeepSeek V3 0324 | 39.0% | 65.4% | estimated ± 7.6 pp, low confidence |
| GPT-6.1 Sol | 54.2% | 84.5% | estimated ± 7.6 pp, low confidence |
| GPT-6 Luna | 54.6% | 85.0% | estimated ± 7.6 pp, low confidence |
| GPT-6 Sol | 57.6% | 88.8% | estimated ± 7.6 pp, low confidence |
| Ling 3.0 Flash VL | 44.2% | 71.9% | estimated ± 7.6 pp, low confidence |
| Ling 3.0 Tiny | 24.2% | 46.7% | estimated ± 7.6 pp, low confidence |
| Ling 3.1 Flash | 54.1% | 84.4% | estimated ± 7.6 pp, low confidence |
| MiMo-V2.6-Flash | 51.3% | 80.9% | estimated ± 7.6 pp, low confidence |
| MiMo-V2.6-Pro | 60.9% | 93.0% | estimated ± 7.6 pp, low confidence |
| North Mini Code | 38.8% | 65.1% | estimated ± 7.6 pp, low confidence |
| Solar Pro 3 | 25.5% | 48.4% | estimated ± 7.6 pp, low confidence |