benchgap
Calibration

AA Coding Index → DeepSWE

DeepSWE is estimated from AA Coding Index with a Hill curve fitted on 20 models measured on both: y = 0.0000 + (1.0062 − 0.0000)·x^6.00 / (0.67629^6.00 + x^6.00), R² = 0.61, cross-validated error 6.4 pp. It is used for 73 estimates.

Estimated modelAA Coding IndexDeepSWESource
Apodex 1.160.8%34.7%estimated ± 6.4 pp, low confidence
Apodex 1.1 Mini60.8%34.7%estimated ± 6.4 pp, low confidence
Celeris-114.4%0.0%estimated ± 6.4 pp, low confidence
Claude 3 Opus19.5%0.1%estimated ± 6.4 pp, low confidence
Claude Opus 4.7 (Adaptive)73.6%62.8%estimated ± 6.4 pp, medium confidence
Command A+27.9%0.5%estimated ± 6.4 pp, low confidence
DeepSeek V323.0%0.2%estimated ± 6.4 pp, low confidence
Gemini 1.5 Pro23.6%0.2%estimated ± 6.4 pp, low confidence
Gemini 2.5 Pro33.3%1.4%estimated ± 6.4 pp, low confidence
Gemini 3.5 Flash70.1%55.8%estimated ± 6.4 pp, medium confidence
Gemini 3.5 Flash-Lite49.3%13.2%estimated ± 6.4 pp, low confidence
Gemma 3 27B10.1%0.0%estimated ± 6.4 pp, low confidence
Gemma 4 12B31.0%0.9%estimated ± 6.4 pp, low confidence
Gemma 4 26B A4B39.3%3.7%estimated ± 6.4 pp, low confidence
Gemma 4 31B43.4%6.6%estimated ± 6.4 pp, low confidence
Gemma 4 E2B7.2%0.0%estimated ± 6.4 pp, low confidence
Gemma 4 E4B9.4%0.0%estimated ± 6.4 pp, low confidence
GLM-4.745.3%8.3%estimated ± 6.4 pp, low confidence
GLM-5.155.8%24.1%estimated ± 6.4 pp, low confidence
GPT-4.1 mini20.2%0.1%estimated ± 6.4 pp, low confidence
GPT-4.1 nano11.1%0.0%estimated ± 6.4 pp, low confidence
GPT-4 Turbo21.5%0.1%estimated ± 6.4 pp, low confidence
GPT-4o mini11.4%0.0%estimated ± 6.4 pp, low confidence
GPT-5.149.4%13.3%estimated ± 6.4 pp, low confidence
GPT-5.4 nano56.1%24.7%estimated ± 6.4 pp, low confidence
GPT-5 (high)37.8%3.0%estimated ± 6.4 pp, low confidence
GPT-OSS 120B30.4%0.8%estimated ± 6.4 pp, low confidence
GPT-OSS 20B20.7%0.1%estimated ± 6.4 pp, low confidence
Granite 4.2 8B22.4%0.1%estimated ± 6.4 pp, low confidence
Grok 4.342.3%5.6%estimated ± 6.4 pp, low confidence
Hy358.8%30.4%estimated ± 6.4 pp, low confidence
Hy3 Preview58.8%30.4%estimated ± 6.4 pp, low confidence
Inkling-Small52.9%18.8%estimated ± 6.4 pp, low confidence
K-Exaone32.1%1.1%estimated ± 6.4 pp, low confidence
Kimi K2.661.8%37.0%estimated ± 6.4 pp, low confidence
Kimi K2.546.8%9.9%estimated ± 6.4 pp, low confidence
Kimi K2.5 (Reasoning)46.8%9.9%estimated ± 6.4 pp, low confidence
Kimi K2.7 Code60.8%34.7%estimated ± 6.4 pp, low confidence
LFM2.5-2.6B7.7%0.0%estimated ± 6.4 pp, low confidence
Ling 2.6 Flash25.3%0.3%estimated ± 6.4 pp, low confidence
Ling 3.0 Flash50.6%15.1%estimated ± 6.4 pp, low confidence
Ling 3.0 Flash FP850.6%15.1%estimated ± 6.4 pp, low confidence
Llama 4 Maverick16.3%0.0%estimated ± 6.4 pp, low confidence
Llama 4 Scout8.2%0.0%estimated ± 6.4 pp, low confidence
MiMo-V2.5-Pro60.2%33.4%estimated ± 6.4 pp, low confidence
MiMo-V2-Flash49.8%13.9%estimated ± 6.4 pp, low confidence
MiniMax M2.752.6%18.3%estimated ± 6.4 pp, low confidence
MiniMax M358.6%29.9%estimated ± 6.4 pp, low confidence
Mistral Large 320.1%0.1%estimated ± 6.4 pp, low confidence
Mistral Medium 3.5 128B46.9%10.1%estimated ± 6.4 pp, low confidence
Mistral Small 426.6%0.4%estimated ± 6.4 pp, low confidence
Mistral Small 4 (Reasoning)26.6%0.4%estimated ± 6.4 pp, low confidence
Muse Glimmer 30B49.0%12.7%estimated ± 6.4 pp, low confidence
Muse Spark58.6%30.0%estimated ± 6.4 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%0.4%estimated ± 6.4 pp, low confidence
Nemotron 3 Nano 30B14.4%0.0%estimated ± 6.4 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%0.0%estimated ± 6.4 pp, low confidence
Nemotron 3 Super 100B37.7%2.9%estimated ± 6.4 pp, low confidence
Nemotron 3 Ultra49.3%13.1%estimated ± 6.4 pp, low confidence
o139.7%4.0%estimated ± 6.4 pp, low confidence
o1-preview34.1%1.6%estimated ± 6.4 pp, low confidence
Quasar 438B61.2%35.7%estimated ± 6.4 pp, low confidence
Qwen3.5-122B-A10B45.7%8.8%estimated ± 6.4 pp, low confidence
Qwen3.6-27B53.7%20.2%estimated ± 6.4 pp, low confidence
Qwen3.6-35B-A3B41.9%5.4%estimated ± 6.4 pp, low confidence
Qwen3.6 Plus54.5%21.7%estimated ± 6.4 pp, low confidence
Qwen3.7 Max66.0%46.6%estimated ± 6.4 pp, low confidence
Qwen3.7 Plus55.9%24.3%estimated ± 6.4 pp, low confidence
Qwen3.8 Max Preview71.8%59.3%estimated ± 6.4 pp, medium confidence
Step 3.7 Flash39.6%3.9%estimated ± 6.4 pp, low confidence
Trinity-Large-Preview25.8%0.3%estimated ± 6.4 pp, low confidence
Trinity-Large-Thinking25.8%0.3%estimated ± 6.4 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%0.0%estimated ± 6.4 pp, low confidence