Calibration
FrontierCode 1.1 Main → Vibe Code Bench
Vibe Code Bench is estimated from FrontierCode 1.1 Main with a Michaelis–Menten curve fitted on 5 models measured on both: y = 1.6468·x / (0.55280 + x), R² = 0.85, cross-validated error 5.6 pp. It is used for 11 estimates.
| Estimated model | FrontierCode 1.1 Main | Vibe Code Bench | Source |
|---|---|---|---|
| Claude Fable 5 | 53.5% | 81.0% | estimated ± 5.6 pp, low confidence |
| Claude Haiku 5.5 | 46.4% | 75.2% | estimated ± 5.6 pp, low confidence |
| Claude Opus 4.8 | 46.5% | 75.2% | estimated ± 5.6 pp, low confidence |
| Claude Opus 5 | 53.4% | 80.9% | estimated ± 5.6 pp, low confidence |
| Claude Opus 5.5 | 54.4% | 81.7% | estimated ± 5.6 pp, low confidence |
| Claude Sonnet 5 | 42.7% | 71.8% | estimated ± 5.6 pp, low confidence |
| Claude Sonnet 5.5 | 46.2% | 75.0% | estimated ± 5.6 pp, low confidence |
| Gemini 3.7 Flash | 43.6% | 72.6% | estimated ± 5.6 pp, low confidence |
| GPT-6 Astra | 53.3% | 80.8% | estimated ± 5.6 pp, low confidence |
| SWE-1.7 | 42.3% | 71.4% | estimated ± 5.6 pp, low confidence |
| SWE-2 | 50.0% | 78.2% | estimated ± 5.6 pp, low confidence |