benchgap
Calibration

SWE-bench Verified → Vibe Code Bench

Vibe Code Bench is estimated from SWE-bench Verified with a offset logistic curve fitted on 11 models measured on both: y = 0.0065 + (0.6246 − 0.0065) / (1 + exp(−83.10·(x − 0.7867))), R² = 0.92, cross-validated error 8.1 pp. It is used for 60 estimates.

Estimated modelSWE-bench VerifiedVibe Code BenchSource
Apodex 1.177.7%19.7%estimated ± 8.1 pp, medium confidence
BTL-478.4%28.1%estimated ± 8.1 pp, medium confidence
Claude 3.5 Sonnet49.0%0.6%estimated ± 8.1 pp, low confidence
Claude 4.1 Opus74.5%2.5%estimated ± 8.1 pp, medium confidence
Claude 4 Sonnet72.7%1.1%estimated ± 8.1 pp, medium confidence
Claude Haiku 4.573.3%1.4%estimated ± 8.1 pp, medium confidence
Claude Mythos 595.5%62.5%estimated ± 8.1 pp, low confidence
Claude Opus 4.580.9%54.1%estimated ± 8.1 pp, medium confidence
Claude Opus 4.7 (Adaptive)87.6%62.4%estimated ± 8.1 pp, low confidence
Claude Sonnet 4.577.2%14.7%estimated ± 8.1 pp, medium confidence
DeepSeek V342.0%0.6%estimated ± 8.1 pp, low confidence
DeepSeek V4 Flash 073179.0%35.7%estimated ± 8.1 pp, medium confidence
dots3-note Preview78.4%28.1%estimated ± 8.1 pp, medium confidence
Ember-192.2%62.5%estimated ± 8.1 pp, low confidence
GLM-4.773.8%1.7%estimated ± 8.1 pp, medium confidence
GPT-4.154.6%0.6%estimated ± 8.1 pp, low confidence
GPT-4.1 mini23.6%0.6%estimated ± 8.1 pp, low confidence
Granite 4.2 30B57.0%0.6%estimated ± 8.1 pp, low confidence
Granite 4.2 8B47.7%0.6%estimated ± 8.1 pp, low confidence
Grok Code Fast 170.8%0.7%estimated ± 8.1 pp, medium confidence
Hy3 Preview74.4%2.4%estimated ± 8.1 pp, medium confidence
Inkling77.6%18.6%estimated ± 8.1 pp, medium confidence
Inkling-Small80.2%48.9%estimated ± 8.1 pp, medium confidence
K-EXAONE 2.068.2%0.7%estimated ± 8.1 pp, medium confidence
Laguna M.174.6%2.7%estimated ± 8.1 pp, medium confidence
Laguna XS.269.9%0.7%estimated ± 8.1 pp, medium confidence
Laguna XS 2.170.9%0.7%estimated ± 8.1 pp, medium confidence
LLaDA2.2-flash49.3%0.6%estimated ± 8.1 pp, low confidence
LongCat-Flash-Lite-Sparse68.2%0.7%estimated ± 8.1 pp, medium confidence
MAI-Code-1.1-Flash72.6%1.0%estimated ± 8.1 pp, medium confidence
MAI-Thinking-173.5%1.5%estimated ± 8.1 pp, medium confidence
MiMo-V2-Flash73.4%1.4%estimated ± 8.1 pp, medium confidence
MiMo-V2-Omni74.8%3.0%estimated ± 8.1 pp, medium confidence
MiMo-V2-Pro78.0%23.1%estimated ± 8.1 pp, medium confidence
MiniCPM5-2B46.4%0.6%estimated ± 8.1 pp, low confidence
MiniMax M380.5%51.4%estimated ± 8.1 pp, medium confidence
Mistral Medium 3.5 128B77.6%18.6%estimated ± 8.1 pp, medium confidence
Muse Glimmer 30B76.0%6.7%estimated ± 8.1 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP452.8%0.6%estimated ± 8.1 pp, low confidence
Nemotron 3 Ultra71.9%0.9%estimated ± 8.1 pp, medium confidence
o3-mini49.3%0.6%estimated ± 8.1 pp, low confidence
Ornith-1.0-35B75.6%5.1%estimated ± 8.1 pp, medium confidence
Ornith-1.0-397B82.4%59.8%estimated ± 8.1 pp, medium confidence
Ornith-1.0-9B69.4%0.7%estimated ± 8.1 pp, medium confidence
Ornith-1.5-35B-A3B79.0%35.7%estimated ± 8.1 pp, medium confidence
Ornith-1.5-397B86.0%62.3%estimated ± 8.1 pp, low confidence
Ornith-1.5-9B70.6%0.7%estimated ± 8.1 pp, medium confidence
Qwen3.5-122B-A10B72.0%0.9%estimated ± 8.1 pp, medium confidence
Qwen3.5-27B72.4%1.0%estimated ± 8.1 pp, medium confidence
Qwen3.5-35B-A3B69.2%0.7%estimated ± 8.1 pp, medium confidence
Qwen3.5 397B76.2%7.7%estimated ± 8.1 pp, medium confidence
Qwen3.6-27B77.2%14.7%estimated ± 8.1 pp, medium confidence
Qwen3.6-35B-A3B73.4%1.4%estimated ± 8.1 pp, medium confidence
Qwen3.7 Max80.4%50.6%estimated ± 8.1 pp, medium confidence
Qwen3.7 Plus77.7%19.7%estimated ± 8.1 pp, medium confidence
Beam80.9%54.1%estimated ± 8.1 pp, medium confidence
Solar Open 270.4%0.7%estimated ± 8.1 pp, medium confidence
Solar Pro 470.6%0.7%estimated ± 8.1 pp, medium confidence
Ternary Bonsai 2 27B60.8%0.6%estimated ± 8.1 pp, low confidence
ZAYA1-74B-Preview53.2%0.6%estimated ± 8.1 pp, low confidence