benchgap
Calibration

Vibe Code Bench → AA-SciCode

AA-SciCode is estimated from Vibe Code Bench with a linear curve fitted on 12 models measured on both: y = 0.1451·x + 0.4574, R² = 0.54, cross-validated error 3.8 pp. It is used for 19 estimates.

Estimated modelVibe Code BenchAA-SciCodeSource
Claude Haiku 4.5 Thinking11.4%47.4%estimated ± 3.8 pp, high confidence
Claude Opus 4.5 Thinking20.6%48.7%estimated ± 3.8 pp, high confidence
Claude Opus 4.6 (Adaptive)53.5%53.5%estimated ± 3.8 pp, high confidence
Claude Sonnet 4.5 Thinking22.6%49.0%estimated ± 3.8 pp, high confidence
DeepSeek V3.2 (Thinking)5.1%46.5%estimated ± 3.8 pp, high confidence
Gemini 3.1 Flash-Lite0.0%45.7%estimated ± 3.8 pp, medium confidence
Gemini 3 Flash20.2%48.7%estimated ± 3.8 pp, high confidence
Gemini 3 Pro14.3%47.8%estimated ± 3.8 pp, high confidence
GLM-4.63.1%46.2%estimated ± 3.8 pp, high confidence
GLM-5 (Reasoning)23.4%49.1%estimated ± 3.8 pp, high confidence
GPT-5.1-Codex13.1%47.6%estimated ± 3.8 pp, high confidence
GPT-5.1-Codex-Max22.2%49.0%estimated ± 3.8 pp, high confidence
GPT-5.2-Codex37.9%51.2%estimated ± 3.8 pp, high confidence
GPT-5 mini14.2%47.8%estimated ± 3.8 pp, high confidence
Grok 4.1 Fast (Reasoning)1.2%45.9%estimated ± 3.8 pp, high confidence
Grok 4 Fast (Reasoning)0.0%45.7%estimated ± 3.8 pp, medium confidence
MiniMax M2.514.9%47.9%estimated ± 3.8 pp, high confidence
Qwen3.5 Plus15.7%48.0%estimated ± 3.8 pp, high confidence
Qwen3 Max3.5%46.2%estimated ± 3.8 pp, high confidence