benchgap
Calibration

Vibe Code Bench → FrontierCode 1.1 Main

FrontierCode 1.1 Main is estimated from Vibe Code Bench with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 0.74004·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.85, cross-validated error 4.3 pp. It is used for 20 estimates.

Estimated modelVibe Code BenchFrontierCode 1.1 MainSource
Claude Haiku 4.5 Thinking11.4%4.5%estimated ± 4.3 pp, low confidence
Claude Opus 4.5 Thinking20.6%8.5%estimated ± 4.3 pp, low confidence
Claude Opus 4.6 (Adaptive)53.5%27.0%estimated ± 4.3 pp, medium confidence
Claude Sonnet 4.5 Thinking22.6%9.4%estimated ± 4.3 pp, low confidence
DeepSeek V3.2 (Thinking)5.1%1.9%estimated ± 4.3 pp, low confidence
Gemini 3.1 Flash-Lite0.0%0.0%estimated ± 4.3 pp, low confidence
Gemini 3 Flash20.2%8.3%estimated ± 4.3 pp, low confidence
Gemini 3 Pro14.3%5.7%estimated ± 4.3 pp, low confidence
GLM-4.63.1%1.2%estimated ± 4.3 pp, low confidence
GLM-5 (Reasoning)23.4%9.8%estimated ± 4.3 pp, low confidence
GPT-5.1-Codex13.1%5.2%estimated ± 4.3 pp, low confidence
GPT-5.1-Codex-Max22.2%9.2%estimated ± 4.3 pp, low confidence
GPT-5.2-Codex37.9%17.3%estimated ± 4.3 pp, low confidence
GPT-5 mini14.2%5.6%estimated ± 4.3 pp, low confidence
Grok 4.1 Fast (Reasoning)1.2%0.4%estimated ± 4.3 pp, low confidence
Grok 4 Fast (Reasoning)0.0%0.0%estimated ± 4.3 pp, low confidence
MiniMax M2.514.9%5.9%estimated ± 4.3 pp, low confidence
Mistral Large 478.4%47.7%estimated ± 4.3 pp, low confidence
Qwen3.5 Plus15.7%6.3%estimated ± 4.3 pp, low confidence
Qwen3 Max3.5%1.3%estimated ± 4.3 pp, low confidence