benchgap
Calibration

AA Coding Index → LiveCodeBench

LiveCodeBench is estimated from AA Coding Index with a inverse Michaelis–Menten curve fitted on 6 models measured on both: y = 3.02770·(x − 0.0077) / (0.0077 + 2.0000 − x), R² = 0.78, cross-validated error 8.9 pp. It is used for 96 estimates.

Estimated modelAA Coding IndexLiveCodeBenchSource
Apodex 1.160.8%100.0%estimated ± 8.9 pp, low confidence
Apodex 1.1 Mini60.8%100.0%estimated ± 8.9 pp, low confidence
Celeris-114.4%22.1%estimated ± 8.9 pp, low confidence
Claude 3 Opus19.5%31.3%estimated ± 8.9 pp, low confidence
Claude Fable 576.5%100.0%estimated ± 8.9 pp, low confidence
Claude Fable 5.181.6%100.0%estimated ± 8.9 pp, low confidence
Claude Opus 4.7 (Adaptive)73.6%100.0%estimated ± 8.9 pp, low confidence
Claude Opus 4.874.3%100.0%estimated ± 8.9 pp, low confidence
Claude Opus 578.0%100.0%estimated ± 8.9 pp, low confidence
Claude Sonnet 571.6%100.0%estimated ± 8.9 pp, low confidence
Command A+27.9%47.4%estimated ± 8.9 pp, low confidence
DeepSeek V4 Flash 073169.1%100.0%estimated ± 8.9 pp, low confidence
DeepSeek V4 Pro 081368.8%100.0%estimated ± 8.9 pp, low confidence
Gemini 1.5 Pro23.6%39.1%estimated ± 8.9 pp, low confidence
Gemini 2.5 Pro33.3%58.7%estimated ± 8.9 pp, low confidence
Gemini 3.1 Pro68.8%100.0%estimated ± 8.9 pp, low confidence
Gemini 3.5 Flash70.1%100.0%estimated ± 8.9 pp, low confidence
Gemini 3.5 Flash-Lite49.3%97.1%estimated ± 8.9 pp, low confidence
Gemini 3.6 Flash69.2%100.0%estimated ± 8.9 pp, low confidence
Gemini 3.7 Flash76.1%100.0%estimated ± 8.9 pp, low confidence
Gemini 3.8 Flash76.3%100.0%estimated ± 8.9 pp, low confidence
Gemma 3 27B10.1%14.7%estimated ± 8.9 pp, low confidence
Gemma 4 12B31.0%53.8%estimated ± 8.9 pp, low confidence
Gemma 4 26B A4B39.3%72.3%estimated ± 8.9 pp, low confidence
Gemma 4 31B43.4%82.1%estimated ± 8.9 pp, low confidence
Gemma 4 E2B7.2%10.1%estimated ± 8.9 pp, low confidence
Gemma 4 E4B9.4%13.6%estimated ± 8.9 pp, low confidence
GLM-5.155.8%100.0%estimated ± 8.9 pp, low confidence
GLM-5.268.8%100.0%estimated ± 8.9 pp, low confidence
GLM-5.374.8%100.0%estimated ± 8.9 pp, low confidence
GPT-4.1 mini20.2%32.6%estimated ± 8.9 pp, low confidence
GPT-4.1 nano11.1%16.6%estimated ± 8.9 pp, low confidence
GPT-4 Turbo21.5%35.0%estimated ± 8.9 pp, low confidence
GPT-4o mini11.4%17.0%estimated ± 8.9 pp, low confidence
GPT-5.149.4%97.2%estimated ± 8.9 pp, low confidence
GPT-5.471.1%100.0%estimated ± 8.9 pp, low confidence
GPT-5.4 mini56.1%100.0%estimated ± 8.9 pp, low confidence
GPT-5.4 nano56.1%100.0%estimated ± 8.9 pp, low confidence
GPT-5.574.9%100.0%estimated ± 8.9 pp, low confidence
GPT-5.6 Luna71.5%100.0%estimated ± 8.9 pp, low confidence
GPT-5.6 Sol77.4%100.0%estimated ± 8.9 pp, low confidence
GPT-5.6 Terra76.7%100.0%estimated ± 8.9 pp, low confidence
GPT-5 (high)37.8%68.7%estimated ± 8.9 pp, low confidence
GPT-6 Astra76.9%100.0%estimated ± 8.9 pp, low confidence
GPT-OSS 120B30.4%52.7%estimated ± 8.9 pp, low confidence
GPT-OSS 20B20.7%33.5%estimated ± 8.9 pp, low confidence
Granite 4.2 8B22.4%36.7%estimated ± 8.9 pp, low confidence
Grok 4.342.3%79.2%estimated ± 8.9 pp, low confidence
Grok 4.572.5%100.0%estimated ± 8.9 pp, low confidence
Grok 4.676.8%100.0%estimated ± 8.9 pp, low confidence
Hy358.8%100.0%estimated ± 8.9 pp, low confidence
Hy3 Preview58.8%100.0%estimated ± 8.9 pp, low confidence
Inkling52.1%100.0%estimated ± 8.9 pp, low confidence
Inkling-Small52.9%100.0%estimated ± 8.9 pp, low confidence
K-Exaone32.1%56.3%estimated ± 8.9 pp, low confidence
Kimi K2.661.8%100.0%estimated ± 8.9 pp, low confidence
Kimi K2.546.8%90.5%estimated ± 8.9 pp, low confidence
Kimi K2.5 (Reasoning)46.8%90.5%estimated ± 8.9 pp, low confidence
Kimi K2.7 Code60.8%100.0%estimated ± 8.9 pp, low confidence
Kimi K376.2%100.0%estimated ± 8.9 pp, low confidence
LFM2.5-2.6B7.7%10.9%estimated ± 8.9 pp, low confidence
Ling 2.6 Flash25.3%42.2%estimated ± 8.9 pp, low confidence
Ling 3.0 Flash50.6%100.0%estimated ± 8.9 pp, low confidence
Ling 3.0 Flash FP850.6%100.0%estimated ± 8.9 pp, low confidence
Llama 4 Maverick16.3%25.5%estimated ± 8.9 pp, low confidence
Llama 4 Scout8.2%11.6%estimated ± 8.9 pp, low confidence
MiMo-V2.5-Pro60.2%100.0%estimated ± 8.9 pp, low confidence
MiMo-V2-Flash49.8%98.4%estimated ± 8.9 pp, low confidence
MiniMax M2.752.6%100.0%estimated ± 8.9 pp, low confidence
MiniMax M358.6%100.0%estimated ± 8.9 pp, low confidence
Mistral Large 320.1%32.3%estimated ± 8.9 pp, low confidence
Mistral Medium 3.5 128B46.9%90.8%estimated ± 8.9 pp, low confidence
Mistral Small 426.6%45.0%estimated ± 8.9 pp, low confidence
Mistral Small 4 (Reasoning)26.6%45.0%estimated ± 8.9 pp, low confidence
Muse Glimmer 30B49.0%96.2%estimated ± 8.9 pp, low confidence
Muse Spark58.6%100.0%estimated ± 8.9 pp, low confidence
Muse Spark 1.171.3%100.0%estimated ± 8.9 pp, low confidence
Muse Spark 1.272.2%100.0%estimated ± 8.9 pp, low confidence
Muse Spark 1.375.8%100.0%estimated ± 8.9 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP426.8%45.2%estimated ± 8.9 pp, low confidence
Nemotron 3 Nano 30B14.4%22.1%estimated ± 8.9 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%21.0%estimated ± 8.9 pp, low confidence
Nemotron 3 Super 100B37.7%68.6%estimated ± 8.9 pp, low confidence
Nemotron 3 Ultra49.3%96.9%estimated ± 8.9 pp, low confidence
o139.7%73.2%estimated ± 8.9 pp, low confidence
o1-preview34.1%60.4%estimated ± 8.9 pp, low confidence
Quasar 438B61.2%100.0%estimated ± 8.9 pp, low confidence
Qwen3.5-122B-A10B45.7%87.7%estimated ± 8.9 pp, low confidence
Qwen3.6 Plus54.5%100.0%estimated ± 8.9 pp, low confidence
Qwen3.8-27B68.1%100.0%estimated ± 8.9 pp, low confidence
Qwen3.8-Flash-Next73.1%100.0%estimated ± 8.9 pp, low confidence
Qwen3.8 Max Preview71.8%100.0%estimated ± 8.9 pp, low confidence
Step 3.7 Flash39.6%72.9%estimated ± 8.9 pp, low confidence
Trinity-Large-Preview25.8%43.3%estimated ± 8.9 pp, low confidence
Trinity-Large-Thinking25.8%43.3%estimated ± 8.9 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%17.9%estimated ± 8.9 pp, low confidence