Calibration
AA Coding Index → FrontierSWE v2
FrontierSWE v2 is estimated from AA Coding Index with a Hill curve fitted on 12 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.88544^6.00 + x^6.00), R² = 0.44, cross-validated error 14.8 pp. It is used for 18 estimates.
| Estimated model | AA Coding Index | FrontierSWE v2 | Source |
|---|---|---|---|
| Claude 3 Opus | 19.5% | 0.0% | estimated ± 14.8 pp, low confidence |
| Gemini 1.5 Pro | 23.6% | 0.0% | estimated ± 14.8 pp, low confidence |
| Gemma 4 12B | 31.0% | 0.2% | estimated ± 14.8 pp, low confidence |
| Gemma 4 E2B | 7.2% | 0.0% | estimated ± 14.8 pp, low confidence |
| GPT-4.1 mini | 20.2% | 0.0% | estimated ± 14.8 pp, low confidence |
| GPT-4.1 nano | 11.1% | 0.0% | estimated ± 14.8 pp, low confidence |
| GPT-4 Turbo | 21.5% | 0.0% | estimated ± 14.8 pp, low confidence |
| GPT-4o mini | 11.4% | 0.0% | estimated ± 14.8 pp, low confidence |
| GPT-5.1 | 49.4% | 3.5% | estimated ± 14.8 pp, low confidence |
| GPT-5 (high) | 37.8% | 0.7% | estimated ± 14.8 pp, low confidence |
| K-Exaone | 32.1% | 0.3% | estimated ± 14.8 pp, low confidence |
| Kimi K2.5 (Reasoning) | 46.8% | 2.6% | estimated ± 14.8 pp, low confidence |
| Ling 2.6 Flash | 25.3% | 0.1% | estimated ± 14.8 pp, low confidence |
| MiMo-V2-Flash | 49.8% | 3.7% | estimated ± 14.8 pp, low confidence |
| Nemotron 3 Nano Omni 30B A3B | 13.8% | 0.0% | estimated ± 14.8 pp, low confidence |
| o1 | 39.7% | 1.0% | estimated ± 14.8 pp, low confidence |
| o1-preview | 34.1% | 0.4% | estimated ± 14.8 pp, low confidence |
| Ultravox v0.6 Llama 3.3 70B | 11.9% | 0.0% | estimated ± 14.8 pp, low confidence |