Calibration
SWE-bench Pro → AA-SciCode
AA-SciCode is estimated from SWE-bench Pro with a linear curve fitted on 37 models measured on both: y = 0.5237·x + 0.1893, R² = 0.81, cross-validated error 3.8 pp. It is used for 14 estimates.
| Estimated model | SWE-bench Pro | AA-SciCode | Source |
|---|---|---|---|
| Atria Dawn Preview | 59.6% | 50.1% | estimated ± 3.8 pp, high confidence |
| Claude Mythos 5 | 80.3% | 61.0% | estimated ± 3.8 pp, high confidence |
| Claude Opus 4.6 | 53.4% | 46.9% | estimated ± 3.8 pp, high confidence |
| GLM-5 | 55.1% | 47.8% | estimated ± 3.8 pp, high confidence |
| GPT-5.2 | 55.6% | 48.1% | estimated ± 3.8 pp, high confidence |
| GPT-5.3 Codex | 56.8% | 48.7% | estimated ± 3.8 pp, high confidence |
| Grok 4.20 | 51.8% | 46.1% | estimated ± 3.8 pp, high confidence |
| Laguna M.1 | 49.2% | 44.7% | estimated ± 3.8 pp, high confidence |
| Laguna S 2.1 | 59.4% | 50.0% | estimated ± 3.8 pp, high confidence |
| Laguna XS.2 | 46.3% | 43.2% | estimated ± 3.8 pp, high confidence |
| Laguna XS 2.1 | 47.6% | 43.9% | estimated ± 3.8 pp, high confidence |
| LLaDA2.2-flash | 30.1% | 34.7% | estimated ± 3.8 pp, high confidence |
| LongCat-Flash-Lite-Sparse | 40.6% | 40.2% | estimated ± 3.8 pp, high confidence |
| MiMo-V2.5 | 56.1% | 48.3% | estimated ± 3.8 pp, high confidence |