benchgap
Calibration

FrontierCode 1.1 Main → Vibe Code Bench

Vibe Code Bench is estimated from FrontierCode 1.1 Main with a Michaelis–Menten curve fitted on 5 models measured on both: y = 1.6468·x / (0.55280 + x), R² = 0.85, cross-validated error 5.6 pp. It is used for 11 estimates.

Estimated modelFrontierCode 1.1 MainVibe Code BenchSource
Claude Fable 553.5%81.0%estimated ± 5.6 pp, low confidence
Claude Haiku 5.546.4%75.2%estimated ± 5.6 pp, low confidence
Claude Opus 4.846.5%75.2%estimated ± 5.6 pp, low confidence
Claude Opus 553.4%80.9%estimated ± 5.6 pp, low confidence
Claude Opus 5.554.4%81.7%estimated ± 5.6 pp, low confidence
Claude Sonnet 542.7%71.8%estimated ± 5.6 pp, low confidence
Claude Sonnet 5.546.2%75.0%estimated ± 5.6 pp, low confidence
Gemini 3.7 Flash43.6%72.6%estimated ± 5.6 pp, low confidence
GPT-6 Astra53.3%80.8%estimated ± 5.6 pp, low confidence
SWE-1.742.3%71.4%estimated ± 5.6 pp, low confidence
SWE-250.0%78.2%estimated ± 5.6 pp, low confidence