benchgap
Calibration

AA Coding Index → SWE-bench Verified

SWE-bench Verified is estimated from AA Coding Index with a Hill curve fitted on 32 models measured on both: y = 0.0000 + (0.9427 − 0.0000)·x^2.17 / (0.25575^2.17 + x^2.17), R² = 0.91, cross-validated error 5.5 pp. It is used for 35 estimates.

Estimated modelAA Coding IndexSWE-bench VerifiedSource
Apodex 1.1 Mini60.8%81.8%estimated ± 5.5 pp, medium confidence
Celeris-114.4%21.0%estimated ± 5.5 pp, low confidence
Claude 3 Opus19.5%33.7%estimated ± 5.5 pp, low confidence
Command A+27.9%51.5%estimated ± 5.5 pp, medium confidence
Gemini 1.5 Pro23.6%43.1%estimated ± 5.5 pp, medium confidence
Gemma 3 27B10.1%11.0%estimated ± 5.5 pp, low confidence
Gemma 4 12B31.0%56.8%estimated ± 5.5 pp, medium confidence
Gemma 4 26B A4B39.3%67.7%estimated ± 5.5 pp, medium confidence
Gemma 4 E2B7.2%5.7%estimated ± 5.5 pp, low confidence
Gemma 4 E4B9.4%9.6%estimated ± 5.5 pp, low confidence
GPT-4.1 nano11.1%13.3%estimated ± 5.5 pp, low confidence
GPT-4 Turbo21.5%38.3%estimated ± 5.5 pp, medium confidence
GPT-4o mini11.4%13.9%estimated ± 5.5 pp, low confidence
GPT-OSS 120B30.4%55.9%estimated ± 5.5 pp, medium confidence
GPT-OSS 20B20.7%36.5%estimated ± 5.5 pp, medium confidence
Hy358.8%81.0%estimated ± 5.5 pp, medium confidence
K-Exaone32.1%58.5%estimated ± 5.5 pp, medium confidence
LFM2.5-2.6B7.7%6.6%estimated ± 5.5 pp, low confidence
Ling 2.6 Flash25.3%46.5%estimated ± 5.5 pp, medium confidence
Ling 3.0 Flash FP850.6%76.8%estimated ± 5.5 pp, medium confidence
Llama 4 Maverick16.3%25.7%estimated ± 5.5 pp, low confidence
Llama 4 Scout8.2%7.3%estimated ± 5.5 pp, low confidence
Mistral Large 320.1%35.0%estimated ± 5.5 pp, low confidence
Mistral Small 426.6%49.2%estimated ± 5.5 pp, medium confidence
Mistral Small 4 (Reasoning)26.6%49.2%estimated ± 5.5 pp, medium confidence
Nemotron 3 Nano 30B14.4%21.0%estimated ± 5.5 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%19.5%estimated ± 5.5 pp, low confidence
Nemotron 3 Super 100B37.7%65.9%estimated ± 5.5 pp, medium confidence
o139.7%68.1%estimated ± 5.5 pp, medium confidence
o1-preview34.1%61.3%estimated ± 5.5 pp, medium confidence
Quasar 438B61.2%81.9%estimated ± 5.5 pp, medium confidence
Qwen3.8 Max Preview71.8%85.2%estimated ± 5.5 pp, medium confidence
Trinity-Large-Preview25.8%47.5%estimated ± 5.5 pp, medium confidence
Trinity-Large-Thinking25.8%47.5%estimated ± 5.5 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B11.9%15.1%estimated ± 5.5 pp, low confidence