benchgap
Calibration

AA Coding Index → NL2Repo

NL2Repo is estimated from AA Coding Index with a linear curve fitted on 13 models measured on both: y = 0.7789·x + -0.0285, R² = 0.73, cross-validated error 5.1 pp. It is used for 51 estimates.

Estimated modelAA Coding IndexNL2RepoSource
Apodex 1.160.8%44.5%estimated ± 5.1 pp, medium confidence
Apodex 1.1 Mini60.8%44.5%estimated ± 5.1 pp, medium confidence
Celeris-114.4%8.4%estimated ± 5.1 pp, low confidence
Claude 3 Opus19.5%12.4%estimated ± 5.1 pp, low confidence
Claude Fable 5.181.6%60.7%estimated ± 5.1 pp, low confidence
Claude Opus 4.7 (Adaptive)73.6%54.5%estimated ± 5.1 pp, medium confidence
Command A+27.9%18.8%estimated ± 5.1 pp, low confidence
DeepSeek V323.0%15.1%estimated ± 5.1 pp, low confidence
Gemini 1.5 Pro23.6%15.6%estimated ± 5.1 pp, low confidence
Gemma 3 27B10.1%5.0%estimated ± 5.1 pp, low confidence
Gemma 4 26B A4B39.3%27.8%estimated ± 5.1 pp, low confidence
Gemma 4 31B43.4%31.0%estimated ± 5.1 pp, medium confidence
Gemma 4 E2B7.2%2.8%estimated ± 5.1 pp, low confidence
Gemma 4 E4B9.4%4.5%estimated ± 5.1 pp, low confidence
GPT-4.1 mini20.2%12.9%estimated ± 5.1 pp, low confidence
GPT-4.1 nano11.1%5.8%estimated ± 5.1 pp, low confidence
GPT-4 Turbo21.5%13.9%estimated ± 5.1 pp, low confidence
GPT-4o mini11.4%6.0%estimated ± 5.1 pp, low confidence
GPT-5.149.4%35.6%estimated ± 5.1 pp, medium confidence
GPT-5.471.1%52.5%estimated ± 5.1 pp, medium confidence
GPT-5 (high)37.8%26.6%estimated ± 5.1 pp, low confidence
GPT-6 Astra76.9%57.1%estimated ± 5.1 pp, low confidence
GPT-OSS 120B30.4%20.9%estimated ± 5.1 pp, low confidence
GPT-OSS 20B20.7%13.3%estimated ± 5.1 pp, low confidence
Hy358.8%43.0%estimated ± 5.1 pp, medium confidence
Hy3 Preview58.8%43.0%estimated ± 5.1 pp, medium confidence
K-Exaone32.1%22.2%estimated ± 5.1 pp, low confidence
Kimi K2.5 (Reasoning)46.8%33.6%estimated ± 5.1 pp, medium confidence
Ling 2.6 Flash25.3%16.8%estimated ± 5.1 pp, low confidence
Ling 3.0 Flash FP850.6%36.6%estimated ± 5.1 pp, medium confidence
Llama 4 Maverick16.3%9.8%estimated ± 5.1 pp, low confidence
Llama 4 Scout8.2%3.5%estimated ± 5.1 pp, low confidence
MiMo-V2-Flash49.8%36.0%estimated ± 5.1 pp, medium confidence
Mistral Large 320.1%12.8%estimated ± 5.1 pp, low confidence
Mistral Small 426.6%17.9%estimated ± 5.1 pp, low confidence
Mistral Small 4 (Reasoning)26.6%17.9%estimated ± 5.1 pp, low confidence
Muse Glimmer 30B49.0%35.3%estimated ± 5.1 pp, medium confidence
Muse Spark 1.375.8%56.2%estimated ± 5.1 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%18.0%estimated ± 5.1 pp, low confidence
Nemotron 3 Nano 30B14.4%8.3%estimated ± 5.1 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%7.9%estimated ± 5.1 pp, low confidence
Nemotron 3 Super 100B37.7%26.5%estimated ± 5.1 pp, low confidence
o139.7%28.1%estimated ± 5.1 pp, low confidence
o1-preview34.1%23.7%estimated ± 5.1 pp, low confidence
Quasar 438B61.2%44.8%estimated ± 5.1 pp, medium confidence
Qwen3.5-122B-A10B45.7%32.8%estimated ± 5.1 pp, medium confidence
Qwen3.8 Max Preview71.8%53.1%estimated ± 5.1 pp, medium confidence
Step 3.7 Flash39.6%28.0%estimated ± 5.1 pp, low confidence
Trinity-Large-Preview25.8%17.2%estimated ± 5.1 pp, low confidence
Trinity-Large-Thinking25.8%17.2%estimated ± 5.1 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%6.4%estimated ± 5.1 pp, low confidence