benchgap
Calibration

AA Coding Index → SWE-Rebench

SWE-Rebench is estimated from AA Coding Index with a offset logistic curve fitted on 5 models measured on both: y = 0.0000 + (0.5809 − 0.0000) / (1 + exp(−200.00·(x − 0.4296))), R² = 0.78, cross-validated error 8.9 pp. It is used for 59 estimates.

Estimated modelAA Coding IndexSWE-RebenchSource
Apodex 1.1 Mini60.8%58.1%estimated ± 8.9 pp, low confidence
Celeris-114.4%0.0%estimated ± 8.9 pp, low confidence
Claude 3 Opus19.5%0.0%estimated ± 8.9 pp, low confidence
Claude Fable 5.181.6%58.1%estimated ± 8.9 pp, low confidence
Command A+27.9%0.0%estimated ± 8.9 pp, low confidence
Gemini 1.5 Pro23.6%0.0%estimated ± 8.9 pp, low confidence
Gemini 3.5 Flash-Lite49.3%58.1%estimated ± 8.9 pp, low confidence
Gemini 3.6 Flash69.2%58.1%estimated ± 8.9 pp, low confidence
Gemini 3.7 Flash76.1%58.1%estimated ± 8.9 pp, low confidence
Gemini 3.8 Flash76.3%58.1%estimated ± 8.9 pp, low confidence
Gemma 3 27B10.1%0.0%estimated ± 8.9 pp, low confidence
Gemma 4 12B31.0%0.0%estimated ± 8.9 pp, low confidence
Gemma 4 26B A4B39.3%0.0%estimated ± 8.9 pp, low confidence
Gemma 4 E2B7.2%0.0%estimated ± 8.9 pp, low confidence
Gemma 4 E4B9.4%0.0%estimated ± 8.9 pp, low confidence
GLM-5.268.8%58.1%estimated ± 8.9 pp, low confidence
GLM-5.374.8%58.1%estimated ± 8.9 pp, low confidence
GPT-4.1 nano11.1%0.0%estimated ± 8.9 pp, low confidence
GPT-4 Turbo21.5%0.0%estimated ± 8.9 pp, low confidence
GPT-4o mini11.4%0.0%estimated ± 8.9 pp, low confidence
GPT-5.6 Luna71.5%58.1%estimated ± 8.9 pp, low confidence
GPT-5.6 Sol77.4%58.1%estimated ± 8.9 pp, low confidence
GPT-5.6 Terra76.7%58.1%estimated ± 8.9 pp, low confidence
GPT-6 Astra76.9%58.1%estimated ± 8.9 pp, low confidence
GPT-OSS 120B30.4%0.0%estimated ± 8.9 pp, low confidence
GPT-OSS 20B20.7%0.0%estimated ± 8.9 pp, low confidence
Grok 4.342.3%11.2%estimated ± 8.9 pp, low confidence
Grok 4.572.5%58.1%estimated ± 8.9 pp, low confidence
Grok 4.676.8%58.1%estimated ± 8.9 pp, low confidence
Hy358.8%58.1%estimated ± 8.9 pp, low confidence
K-Exaone32.1%0.0%estimated ± 8.9 pp, low confidence
Kimi K2.7 Code60.8%58.1%estimated ± 8.9 pp, low confidence
Kimi K376.2%58.1%estimated ± 8.9 pp, low confidence
LFM2.5-2.6B7.7%0.0%estimated ± 8.9 pp, low confidence
Ling 2.6 Flash25.3%0.0%estimated ± 8.9 pp, low confidence
Ling 3.0 Flash50.6%58.1%estimated ± 8.9 pp, low confidence
Ling 3.0 Flash FP850.6%58.1%estimated ± 8.9 pp, low confidence
Llama 4 Maverick16.3%0.0%estimated ± 8.9 pp, low confidence
Llama 4 Scout8.2%0.0%estimated ± 8.9 pp, low confidence
MiMo-V2.5-Pro60.2%58.1%estimated ± 8.9 pp, low confidence
Mistral Large 320.1%0.0%estimated ± 8.9 pp, low confidence
Mistral Small 426.6%0.0%estimated ± 8.9 pp, low confidence
Mistral Small 4 (Reasoning)26.6%0.0%estimated ± 8.9 pp, low confidence
Muse Spark 1.171.3%58.1%estimated ± 8.9 pp, low confidence
Muse Spark 1.272.2%58.1%estimated ± 8.9 pp, low confidence
Muse Spark 1.375.8%58.1%estimated ± 8.9 pp, low confidence
Nemotron 3 Nano 30B14.4%0.0%estimated ± 8.9 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%0.0%estimated ± 8.9 pp, low confidence
Nemotron 3 Super 100B37.7%0.0%estimated ± 8.9 pp, low confidence
o139.7%0.1%estimated ± 8.9 pp, low confidence
o1-preview34.1%0.0%estimated ± 8.9 pp, low confidence
Quasar 438B61.2%58.1%estimated ± 8.9 pp, low confidence
Qwen3.8-27B68.1%58.1%estimated ± 8.9 pp, low confidence
Qwen3.8-Flash-Next73.1%58.1%estimated ± 8.9 pp, low confidence
Qwen3.8 Max Preview71.8%58.1%estimated ± 8.9 pp, low confidence
Step 3.7 Flash39.6%0.1%estimated ± 8.9 pp, low confidence
Trinity-Large-Preview25.8%0.0%estimated ± 8.9 pp, low confidence
Trinity-Large-Thinking25.8%0.0%estimated ± 8.9 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%0.0%estimated ± 8.9 pp, low confidence