benchgap
Calibration

DeepSWE → AA Coding Index

AA Coding Index is estimated from DeepSWE with a linear curve fitted on 20 models measured on both: y = 0.3085·x + 0.5457, R² = 0.59, cross-validated error 2.5 pp. It is used for 20 estimates.

Estimated modelDeepSWEAA Coding IndexSource
DeepSeek V4.1 Flash74.2%77.5%estimated ± 2.5 pp, high confidence
Ember-175.2%77.8%estimated ± 2.5 pp, high confidence
Gemini 4 Argon77.9%78.6%estimated ± 2.5 pp, medium confidence
GLM-5.3-Flash63.4%74.1%estimated ± 2.5 pp, high confidence
GPT-6.1 Sol71.9%76.7%estimated ± 2.5 pp, high confidence
GPT-6 Luna66.6%75.1%estimated ± 2.5 pp, high confidence
GPT-6 Sol68.8%75.8%estimated ± 2.5 pp, high confidence
Hy4 preview64.3%74.4%estimated ± 2.5 pp, high confidence
Laguna S 2.140.4%67.0%estimated ± 2.5 pp, medium confidence
MiMo-V2.6-Flash67.9%75.5%estimated ± 2.5 pp, high confidence
MiMo-V2.6-Pro71.9%76.7%estimated ± 2.5 pp, high confidence
Ornith-1.5-35B-A3B22.0%61.4%estimated ± 2.5 pp, medium confidence
Ornith-1.5-397B56.0%71.8%estimated ± 2.5 pp, high confidence
Pareto 26.10 Preview69.9%76.1%estimated ± 2.5 pp, high confidence
Pareto 26.974.0%77.4%estimated ± 2.5 pp, high confidence
Qwen3.8 Max56.6%72.0%estimated ± 2.5 pp, high confidence
Qwen3.8-Omni-Flash57.8%72.4%estimated ± 2.5 pp, high confidence
Beam44.4%68.3%estimated ± 2.5 pp, high confidence
Step 5 Preview67.7%75.5%estimated ± 2.5 pp, high confidence
SWE-273.0%77.1%estimated ± 2.5 pp, high confidence