benchgap
Calibration

AA Coding Index → SWE-bench Pro

SWE-bench Pro is estimated from AA Coding Index with a Michaelis–Menten curve fitted on 38 models measured on both: y = 1.3899·x / (0.83100 + x), R² = 0.68, cross-validated error 6.0 pp. It is used for 36 estimates.

Estimated modelAA Coding IndexSWE-bench ProSource
Apodex 1.1 Mini60.8%58.7%estimated ± 6.0 pp, medium confidence
Celeris-114.4%20.5%estimated ± 6.0 pp, low confidence
Claude 3 Opus19.5%26.4%estimated ± 6.0 pp, low confidence
Command A+27.9%34.9%estimated ± 6.0 pp, medium confidence
Gemini 1.5 Pro23.6%30.8%estimated ± 6.0 pp, medium confidence
Gemma 3 27B10.1%15.0%estimated ± 6.0 pp, low confidence
Gemma 4 26B A4B39.3%44.6%estimated ± 6.0 pp, medium confidence
Gemma 4 31B43.4%47.7%estimated ± 6.0 pp, medium confidence
Gemma 4 E2B7.2%11.1%estimated ± 6.0 pp, low confidence
Gemma 4 E4B9.4%14.1%estimated ± 6.0 pp, low confidence
GPT-4.1 nano11.1%16.4%estimated ± 6.0 pp, low confidence
GPT-4 Turbo21.5%28.6%estimated ± 6.0 pp, low confidence
GPT-4o mini11.4%16.7%estimated ± 6.0 pp, low confidence
GPT-5.149.4%51.8%estimated ± 6.0 pp, medium confidence
GPT-5 (high)37.8%43.4%estimated ± 6.0 pp, medium confidence
GPT-OSS 120B30.4%37.3%estimated ± 6.0 pp, medium confidence
GPT-OSS 20B20.7%27.7%estimated ± 6.0 pp, low confidence
Hy358.8%57.6%estimated ± 6.0 pp, medium confidence
K-Exaone32.1%38.7%estimated ± 6.0 pp, medium confidence
Ling 2.6 Flash25.3%32.4%estimated ± 6.0 pp, medium confidence
Ling 3.0 Flash FP850.6%52.6%estimated ± 6.0 pp, medium confidence
Llama 4 Maverick16.3%22.8%estimated ± 6.0 pp, low confidence
Llama 4 Scout8.2%12.4%estimated ± 6.0 pp, low confidence
Mistral Large 320.1%27.0%estimated ± 6.0 pp, low confidence
Mistral Small 426.6%33.7%estimated ± 6.0 pp, medium confidence
Mistral Small 4 (Reasoning)26.6%33.7%estimated ± 6.0 pp, medium confidence
Nemotron 3 Nano 30B14.4%20.5%estimated ± 6.0 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%19.7%estimated ± 6.0 pp, low confidence
Nemotron 3 Super 100B37.7%43.4%estimated ± 6.0 pp, medium confidence
o139.7%44.9%estimated ± 6.0 pp, medium confidence
o1-preview34.1%40.4%estimated ± 6.0 pp, medium confidence
Quasar 438B61.2%58.9%estimated ± 6.0 pp, medium confidence
Qwen3.8 Max Preview71.8%64.4%estimated ± 6.0 pp, medium confidence
Trinity-Large-Preview25.8%32.9%estimated ± 6.0 pp, medium confidence
Trinity-Large-Thinking25.8%32.9%estimated ± 6.0 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B11.9%17.4%estimated ± 6.0 pp, low confidence