benchgap
Calibration

Vibe Code Bench → SWE-bench Verified

SWE-bench Verified is estimated from Vibe Code Bench with a linear curve fitted on 11 models measured on both: y = 0.1886·x + 0.7165, R² = 0.61, cross-validated error 4.4 pp. It is used for 25 estimates.

Estimated modelVibe Code BenchSWE-bench VerifiedSource
Claude Haiku 4.5 Thinking11.4%73.8%estimated ± 4.4 pp, high confidence
Claude Opus 4.5 Thinking20.6%75.5%estimated ± 4.4 pp, high confidence
Claude Opus 4.6 (Adaptive)53.5%81.7%estimated ± 4.4 pp, high confidence
Claude Sonnet 4.5 Thinking22.6%75.9%estimated ± 4.4 pp, high confidence
DeepSeek V3.2 (Thinking)5.1%72.6%estimated ± 4.4 pp, high confidence
Gemini 3.1 Flash-Lite0.0%71.7%estimated ± 4.4 pp, medium confidence
Gemini 3.1 Pro32.0%77.7%estimated ± 4.4 pp, high confidence
Gemini 3 Flash20.2%75.5%estimated ± 4.4 pp, high confidence
Gemini 3 Pro14.3%74.4%estimated ± 4.4 pp, high confidence
Gemini 4 Argon91.9%89.0%estimated ± 4.4 pp, medium confidence
GLM-4.63.1%72.2%estimated ± 4.4 pp, high confidence
GLM-5 (Reasoning)23.4%76.1%estimated ± 4.4 pp, high confidence
GPT-5.124.6%76.3%estimated ± 4.4 pp, high confidence
GPT-5.1-Codex13.1%74.1%estimated ± 4.4 pp, high confidence
GPT-5.1-Codex-Max22.2%75.8%estimated ± 4.4 pp, high confidence
GPT-5.2-Codex37.9%78.8%estimated ± 4.4 pp, high confidence
GPT-5.4 nano26.1%76.6%estimated ± 4.4 pp, high confidence
GPT-5 (high)20.1%75.4%estimated ± 4.4 pp, high confidence
GPT-5 mini14.2%74.3%estimated ± 4.4 pp, high confidence
Grok 4.1 Fast (Reasoning)1.2%71.9%estimated ± 4.4 pp, high confidence
Grok 4 Fast (Reasoning)0.0%71.7%estimated ± 4.4 pp, medium confidence
MiniMax M2.514.9%74.5%estimated ± 4.4 pp, high confidence
Mistral Large 478.4%86.4%estimated ± 4.4 pp, medium confidence
Qwen3.5 Plus15.7%74.6%estimated ± 4.4 pp, high confidence
Qwen3 Max3.5%72.3%estimated ± 4.4 pp, high confidence