benchgap
Calibration

AA Coding Index → Vibe Code Bench

Vibe Code Bench is estimated from AA Coding Index with a inverse Michaelis–Menten curve fitted on 16 models measured on both: y = 0.22036·(x − 0.0000) / (0.0000 + 0.9934 − x), R² = 0.77, cross-validated error 10.3 pp. It is used for 55 estimates.

Estimated modelAA Coding IndexVibe Code BenchSource
Apodex 1.1 Mini60.8%34.7%estimated ± 10.3 pp, low confidence
Celeris-114.4%3.7%estimated ± 10.3 pp, low confidence
Claude 3 Opus19.5%5.4%estimated ± 10.3 pp, low confidence
Claude Fable 5.181.6%100.0%estimated ± 10.3 pp, low confidence
Command A+27.9%8.6%estimated ± 10.3 pp, low confidence
Gemini 1.5 Pro23.6%6.9%estimated ± 10.3 pp, low confidence
Gemini 3.5 Flash-Lite49.3%21.7%estimated ± 10.3 pp, low confidence
Gemini 3.6 Flash69.2%50.7%estimated ± 10.3 pp, low confidence
Gemini 3.8 Flash76.3%72.9%estimated ± 10.3 pp, low confidence
Gemma 3 27B10.1%2.5%estimated ± 10.3 pp, low confidence
Gemma 4 12B31.0%10.0%estimated ± 10.3 pp, low confidence
Gemma 4 26B A4B39.3%14.4%estimated ± 10.3 pp, low confidence
Gemma 4 E2B7.2%1.7%estimated ± 10.3 pp, low confidence
Gemma 4 E4B9.4%2.3%estimated ± 10.3 pp, low confidence
GLM-5.268.8%49.5%estimated ± 10.3 pp, low confidence
GLM-5.374.8%67.0%estimated ± 10.3 pp, low confidence
GPT-4.1 nano11.1%2.8%estimated ± 10.3 pp, low confidence
GPT-4 Turbo21.5%6.1%estimated ± 10.3 pp, low confidence
GPT-4o mini11.4%2.9%estimated ± 10.3 pp, low confidence
GPT-5.6 Luna71.5%56.5%estimated ± 10.3 pp, low confidence
GPT-5.6 Sol77.4%77.7%estimated ± 10.3 pp, low confidence
GPT-5.6 Terra76.7%74.5%estimated ± 10.3 pp, low confidence
Grok 4.342.3%16.3%estimated ± 10.3 pp, low confidence
Grok 4.572.5%59.4%estimated ± 10.3 pp, low confidence
Grok 4.676.8%75.0%estimated ± 10.3 pp, low confidence
Hy358.8%32.0%estimated ± 10.3 pp, low confidence
K-Exaone32.1%10.5%estimated ± 10.3 pp, low confidence
Kimi K2.7 Code60.8%34.7%estimated ± 10.3 pp, low confidence
Kimi K376.2%72.7%estimated ± 10.3 pp, low confidence
LFM2.5-2.6B7.7%1.9%estimated ± 10.3 pp, low confidence
Ling 2.6 Flash25.3%7.5%estimated ± 10.3 pp, low confidence
Ling 3.0 Flash50.6%22.9%estimated ± 10.3 pp, low confidence
Ling 3.0 Flash FP850.6%22.9%estimated ± 10.3 pp, low confidence
Llama 4 Maverick16.3%4.3%estimated ± 10.3 pp, low confidence
Llama 4 Scout8.2%2.0%estimated ± 10.3 pp, low confidence
MiMo-V2.5-Pro60.2%33.9%estimated ± 10.3 pp, low confidence
Mistral Large 320.1%5.6%estimated ± 10.3 pp, low confidence
Mistral Small 426.6%8.1%estimated ± 10.3 pp, low confidence
Mistral Small 4 (Reasoning)26.6%8.1%estimated ± 10.3 pp, low confidence
Muse Spark 1.171.3%56.1%estimated ± 10.3 pp, low confidence
Muse Spark 1.272.2%58.7%estimated ± 10.3 pp, low confidence
Muse Spark 1.375.8%70.9%estimated ± 10.3 pp, low confidence
Nemotron 3 Nano 30B14.4%3.7%estimated ± 10.3 pp, low confidence
Nemotron 3 Nano Omni 30B A3B13.8%3.5%estimated ± 10.3 pp, low confidence
Nemotron 3 Super 100B37.7%13.5%estimated ± 10.3 pp, low confidence
o139.7%14.7%estimated ± 10.3 pp, low confidence
o1-preview34.1%11.5%estimated ± 10.3 pp, low confidence
Quasar 438B61.2%35.4%estimated ± 10.3 pp, low confidence
Qwen3.8-27B68.1%48.0%estimated ± 10.3 pp, low confidence
Qwen3.8-Flash-Next73.1%61.2%estimated ± 10.3 pp, low confidence
Qwen3.8 Max Preview71.8%57.5%estimated ± 10.3 pp, low confidence
Step 3.7 Flash39.6%14.6%estimated ± 10.3 pp, low confidence
Trinity-Large-Preview25.8%7.7%estimated ± 10.3 pp, low confidence
Trinity-Large-Thinking25.8%7.7%estimated ± 10.3 pp, low confidence
Ultravox v0.6 Llama 3.3 70B11.9%3.0%estimated ± 10.3 pp, low confidence