Calibration
DeepSWE → AA Coding Index
AA Coding Index is estimated from DeepSWE with a linear curve fitted on 20 models measured on both: y = 0.3085·x + 0.5457, R² = 0.59, cross-validated error 2.5 pp. It is used for 20 estimates.
| Estimated model | DeepSWE | AA Coding Index | Source |
|---|---|---|---|
| DeepSeek V4.1 Flash | 74.2% | 77.5% | estimated ± 2.5 pp, high confidence |
| Ember-1 | 75.2% | 77.8% | estimated ± 2.5 pp, high confidence |
| Gemini 4 Argon | 77.9% | 78.6% | estimated ± 2.5 pp, medium confidence |
| GLM-5.3-Flash | 63.4% | 74.1% | estimated ± 2.5 pp, high confidence |
| GPT-6.1 Sol | 71.9% | 76.7% | estimated ± 2.5 pp, high confidence |
| GPT-6 Luna | 66.6% | 75.1% | estimated ± 2.5 pp, high confidence |
| GPT-6 Sol | 68.8% | 75.8% | estimated ± 2.5 pp, high confidence |
| Hy4 preview | 64.3% | 74.4% | estimated ± 2.5 pp, high confidence |
| Laguna S 2.1 | 40.4% | 67.0% | estimated ± 2.5 pp, medium confidence |
| MiMo-V2.6-Flash | 67.9% | 75.5% | estimated ± 2.5 pp, high confidence |
| MiMo-V2.6-Pro | 71.9% | 76.7% | estimated ± 2.5 pp, high confidence |
| Ornith-1.5-35B-A3B | 22.0% | 61.4% | estimated ± 2.5 pp, medium confidence |
| Ornith-1.5-397B | 56.0% | 71.8% | estimated ± 2.5 pp, high confidence |
| Pareto 26.10 Preview | 69.9% | 76.1% | estimated ± 2.5 pp, high confidence |
| Pareto 26.9 | 74.0% | 77.4% | estimated ± 2.5 pp, high confidence |
| Qwen3.8 Max | 56.6% | 72.0% | estimated ± 2.5 pp, high confidence |
| Qwen3.8-Omni-Flash | 57.8% | 72.4% | estimated ± 2.5 pp, high confidence |
| Beam | 44.4% | 68.3% | estimated ± 2.5 pp, high confidence |
| Step 5 Preview | 67.7% | 75.5% | estimated ± 2.5 pp, high confidence |
| SWE-2 | 73.0% | 77.1% | estimated ± 2.5 pp, high confidence |