benchgap
Calibration

AA Coding Index → FrontierSWE v2

FrontierSWE v2 is estimated from AA Coding Index with a Hill curve fitted on 12 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.88544^6.00 + x^6.00), R² = 0.44, cross-validated error 14.8 pp. It is used for 18 estimates.

Estimated modelAA Coding IndexFrontierSWE v2Source
Claude 3 Opus19.5%0.0%estimated ± 14.8 pp, low confidence
Gemini 1.5 Pro23.6%0.0%estimated ± 14.8 pp, low confidence
Gemma 4 12B31.0%0.2%estimated ± 14.8 pp, low confidence
Gemma 4 E2B7.2%0.0%estimated ± 14.8 pp, low confidence
GPT-4.1 mini20.2%0.0%estimated ± 14.8 pp, low confidence
GPT-4.1 nano11.1%0.0%estimated ± 14.8 pp, low confidence
GPT-4 Turbo21.5%0.0%estimated ± 14.8 pp, low confidence
GPT-4o mini11.4%0.0%estimated ± 14.8 pp, low confidence
GPT-5.149.4%3.5%estimated ± 14.8 pp, low confidence
GPT-5 (high)37.8%0.7%estimated ± 14.8 pp, low confidence
K-Exaone32.1%0.3%estimated ± 14.8 pp, low confidence
Kimi K2.5 (Reasoning)46.8%2.6%estimated ± 14.8 pp, low confidence
Ling 2.6 Flash25.3%0.1%estimated ± 14.8 pp, low confidence
MiMo-V2-Flash49.8%3.7%estimated ± 14.8 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%0.0%estimated ± 14.8 pp, low confidence
o139.7%1.0%estimated ± 14.8 pp, low confidence
o1-preview34.1%0.4%estimated ± 14.8 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%0.0%estimated ± 14.8 pp, low confidence