benchgap
Calibration

Vibe Code Bench → AA Coding Index

AA Coding Index is estimated from Vibe Code Bench with a linear curve fitted on 16 models measured on both: y = 0.5460·x + 0.3864, R² = 0.73, cross-validated error 6.8 pp. It is used for 16 estimates.

Estimated modelVibe Code BenchAA Coding IndexSource
Claude Haiku 4.5 Thinking11.4%44.9%estimated ± 6.8 pp, medium confidence
Claude Opus 4.5 Thinking20.6%49.9%estimated ± 6.8 pp, medium confidence
Claude Opus 4.6 (Adaptive)53.5%67.8%estimated ± 6.8 pp, medium confidence
Claude Sonnet 4.5 Thinking22.6%51.0%estimated ± 6.8 pp, medium confidence
DeepSeek V3.2 (Thinking)5.1%41.4%estimated ± 6.8 pp, medium confidence
Gemini 3 Pro14.3%46.4%estimated ± 6.8 pp, medium confidence
GLM-4.63.1%40.3%estimated ± 6.8 pp, medium confidence
GLM-5 (Reasoning)23.4%51.4%estimated ± 6.8 pp, medium confidence
GPT-5.1-Codex13.1%45.8%estimated ± 6.8 pp, medium confidence
GPT-5.1-Codex-Max22.2%50.7%estimated ± 6.8 pp, medium confidence
GPT-5 mini14.2%46.4%estimated ± 6.8 pp, medium confidence
Grok 4.1 Fast (Reasoning)1.2%39.3%estimated ± 6.8 pp, medium confidence
Grok 4 Fast (Reasoning)0.0%38.6%estimated ± 6.8 pp, low confidence
MiniMax M2.514.9%46.7%estimated ± 6.8 pp, medium confidence
Qwen3.5 Plus15.7%47.2%estimated ± 6.8 pp, medium confidence
Qwen3 Max3.5%40.5%estimated ± 6.8 pp, medium confidence