benchgap
Calibration

AA Coding Index → LiveCodeBench v6

LiveCodeBench v6 is estimated from AA Coding Index with a Michaelis–Menten + offset curve fitted on 9 models measured on both: y = 0.5038 + 0.7610·x / (0.58411 + x), R² = 0.96, cross-validated error 3.1 pp. It is used for 61 estimates.

Estimated modelAA Coding IndexLiveCodeBench v6Source
Apodex 1.160.8%89.2%estimated ± 3.1 pp, high confidence
Apodex 1.1 Mini60.8%89.2%estimated ± 3.1 pp, high confidence
Celeris-114.4%65.4%estimated ± 3.1 pp, high confidence
Claude 3 Opus19.5%69.5%estimated ± 3.1 pp, high confidence
Command A+27.9%75.0%estimated ± 3.1 pp, high confidence
DeepSeek V323.0%71.9%estimated ± 3.1 pp, high confidence
Gemini 1.5 Pro23.6%72.3%estimated ± 3.1 pp, high confidence
Gemini 2.5 Pro33.3%78.0%estimated ± 3.1 pp, high confidence
Gemini 3.1 Pro68.8%91.5%estimated ± 3.1 pp, high confidence
Gemini 3.6 Flash69.2%91.7%estimated ± 3.1 pp, high confidence
Gemini 3.7 Flash76.1%93.4%estimated ± 3.1 pp, medium confidence
Gemini 3.8 Flash76.3%93.5%estimated ± 3.1 pp, medium confidence
Gemma 3 27B10.1%61.6%estimated ± 3.1 pp, high confidence
Gemma 4 26B A4B39.3%81.0%estimated ± 3.1 pp, high confidence
Gemma 4 31B43.4%82.8%estimated ± 3.1 pp, high confidence
Gemma 4 E2B7.2%58.8%estimated ± 3.1 pp, medium confidence
Gemma 4 E4B9.4%60.9%estimated ± 3.1 pp, high confidence
GLM-4.745.3%83.6%estimated ± 3.1 pp, high confidence
GLM-5.374.8%93.1%estimated ± 3.1 pp, medium confidence
GPT-4.1 mini20.2%69.9%estimated ± 3.1 pp, high confidence
GPT-4.1 nano11.1%62.6%estimated ± 3.1 pp, high confidence
GPT-4 Turbo21.5%70.8%estimated ± 3.1 pp, high confidence
GPT-4o mini11.4%62.8%estimated ± 3.1 pp, high confidence
GPT-5.149.4%85.2%estimated ± 3.1 pp, high confidence
GPT-5.4 mini56.1%87.7%estimated ± 3.1 pp, high confidence
GPT-5.4 nano56.1%87.7%estimated ± 3.1 pp, high confidence
GPT-5 (high)37.8%80.3%estimated ± 3.1 pp, high confidence
GPT-6 Astra76.9%93.6%estimated ± 3.1 pp, medium confidence
GPT-OSS 120B30.4%76.5%estimated ± 3.1 pp, high confidence
GPT-OSS 20B20.7%70.3%estimated ± 3.1 pp, high confidence
Grok 4.342.3%82.3%estimated ± 3.1 pp, high confidence
Grok 4.676.8%93.6%estimated ± 3.1 pp, medium confidence
Hy358.8%88.6%estimated ± 3.1 pp, high confidence
Hy3 Preview58.8%88.6%estimated ± 3.1 pp, high confidence
K-Exaone32.1%77.4%estimated ± 3.1 pp, high confidence
Kimi K2.5 (Reasoning)46.8%84.2%estimated ± 3.1 pp, high confidence
Kimi K2.7 Code60.8%89.2%estimated ± 3.1 pp, high confidence
Kimi K376.2%93.5%estimated ± 3.1 pp, medium confidence
Ling 2.6 Flash25.3%73.4%estimated ± 3.1 pp, high confidence
Ling 3.0 Flash FP850.6%85.7%estimated ± 3.1 pp, high confidence
Llama 4 Maverick16.3%67.0%estimated ± 3.1 pp, high confidence
Llama 4 Scout8.2%59.7%estimated ± 3.1 pp, high confidence
MiMo-V2-Flash49.8%85.4%estimated ± 3.1 pp, high confidence
Mistral Large 320.1%69.8%estimated ± 3.1 pp, high confidence
Mistral Medium 3.5 128B46.9%84.3%estimated ± 3.1 pp, high confidence
Mistral Small 426.6%74.2%estimated ± 3.1 pp, high confidence
Mistral Small 4 (Reasoning)26.6%74.2%estimated ± 3.1 pp, high confidence
Muse Spark 1.272.2%92.5%estimated ± 3.1 pp, high confidence
Muse Spark 1.375.8%93.4%estimated ± 3.1 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%74.3%estimated ± 3.1 pp, high confidence
Nemotron 3 Nano 30B14.4%65.4%estimated ± 3.1 pp, high confidence
Nemotron 3 Nano Omni 30B A3B13.8%64.9%estimated ± 3.1 pp, high confidence
Nemotron 3 Super 100B37.7%80.2%estimated ± 3.1 pp, high confidence
o139.7%81.2%estimated ± 3.1 pp, high confidence
o1-preview34.1%78.4%estimated ± 3.1 pp, high confidence
Quasar 438B61.2%89.3%estimated ± 3.1 pp, high confidence
Qwen3.5-122B-A10B45.7%83.8%estimated ± 3.1 pp, high confidence
Qwen3.8 Max Preview71.8%92.3%estimated ± 3.1 pp, high confidence
Trinity-Large-Preview25.8%73.7%estimated ± 3.1 pp, high confidence
Trinity-Large-Thinking25.8%73.7%estimated ± 3.1 pp, high confidence
Ultravox v0.6 Llama 3.3 70B11.9%63.3%estimated ± 3.1 pp, high confidence