benchgap
Calibration

AA Coding Index → Vals SWE-bench

Vals SWE-bench is estimated from AA Coding Index with a linear curve fitted on 44 models measured on both: y = 0.7381·x + 0.3338, R² = 0.70, cross-validated error 5.4 pp. It is used for 41 estimates.

Estimated modelAA Coding IndexVals SWE-benchSource
Apodex 1.160.8%78.2%estimated ± 5.4 pp, medium confidence
Apodex 1.1 Mini60.8%78.2%estimated ± 5.4 pp, medium confidence
Celeris-114.4%44.0%estimated ± 5.4 pp, low confidence
Claude 3 Opus19.5%47.8%estimated ± 5.4 pp, low confidence
Claude Opus 4.7 (Adaptive)73.6%87.7%estimated ± 5.4 pp, medium confidence
Command A+27.9%53.9%estimated ± 5.4 pp, low confidence
DeepSeek V323.0%50.4%estimated ± 5.4 pp, low confidence
Gemini 1.5 Pro23.6%50.8%estimated ± 5.4 pp, low confidence
Gemma 3 27B10.1%40.8%estimated ± 5.4 pp, low confidence
Gemma 4 12B31.0%56.2%estimated ± 5.4 pp, low confidence
Gemma 4 26B A4B39.3%62.4%estimated ± 5.4 pp, medium confidence
Gemma 4 E2B7.2%38.7%estimated ± 5.4 pp, low confidence
Gemma 4 E4B9.4%40.3%estimated ± 5.4 pp, low confidence
GPT-4.1 mini20.2%48.3%estimated ± 5.4 pp, low confidence
GPT-4.1 nano11.1%41.6%estimated ± 5.4 pp, low confidence
GPT-4 Turbo21.5%49.2%estimated ± 5.4 pp, low confidence
GPT-4o mini11.4%41.8%estimated ± 5.4 pp, low confidence
GPT-5.149.4%69.8%estimated ± 5.4 pp, medium confidence
GPT-5 (high)37.8%61.3%estimated ± 5.4 pp, medium confidence
Hy358.8%76.8%estimated ± 5.4 pp, medium confidence
K-Exaone32.1%57.1%estimated ± 5.4 pp, low confidence
Kimi K2.5 (Reasoning)46.8%67.9%estimated ± 5.4 pp, medium confidence
LFM2.5-2.6B7.7%39.1%estimated ± 5.4 pp, low confidence
Llama 4 Maverick16.3%45.4%estimated ± 5.4 pp, low confidence
Llama 4 Scout8.2%39.4%estimated ± 5.4 pp, low confidence
MiMo-V2-Flash49.8%70.2%estimated ± 5.4 pp, medium confidence
Mistral Large 320.1%48.2%estimated ± 5.4 pp, low confidence
Mistral Small 426.6%53.0%estimated ± 5.4 pp, low confidence
Mistral Small 4 (Reasoning)26.6%53.0%estimated ± 5.4 pp, low confidence
Muse Spark 1.375.8%89.3%estimated ± 5.4 pp, medium confidence
Nemotron 3 Nano 30B14.4%44.0%estimated ± 5.4 pp, low confidence
Nemotron 3 Super 100B37.7%61.2%estimated ± 5.4 pp, medium confidence
o139.7%62.7%estimated ± 5.4 pp, medium confidence
o1-preview34.1%58.5%estimated ± 5.4 pp, medium confidence
Quasar 438B61.2%78.5%estimated ± 5.4 pp, medium confidence
Qwen3.5-122B-A10B45.7%67.1%estimated ± 5.4 pp, medium confidence
Qwen3.8 Max Preview71.8%86.4%estimated ± 5.4 pp, medium confidence
Step 3.7 Flash39.6%62.6%estimated ± 5.4 pp, medium confidence
Trinity-Large-Preview25.8%52.4%estimated ± 5.4 pp, low confidence
Trinity-Large-Thinking25.8%52.4%estimated ± 5.4 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%42.2%estimated ± 5.4 pp, low confidence