Calibration
Vibe Code Bench → Vals SWE-bench
Vals SWE-bench is estimated from Vibe Code Bench with a offset logistic curve fitted on 19 models measured on both: y = 0.0000 + (0.8323 − 0.0000) / (1 + exp(−4.17·(x − -0.2546))), R² = 0.57, cross-validated error 6.3 pp. It is used for 18 estimates.
| Estimated model | Vibe Code Bench | Vals SWE-bench | Source |
|---|---|---|---|
| Claude Haiku 4.5 Thinking | 11.4% | 68.5% | estimated ± 6.3 pp, medium confidence |
| Claude Opus 4.5 Thinking | 20.6% | 72.6% | estimated ± 6.3 pp, medium confidence |
| Claude Opus 4.6 (Adaptive) | 53.5% | 80.2% | estimated ± 6.3 pp, medium confidence |
| Claude Sonnet 4.5 Thinking | 22.6% | 73.3% | estimated ± 6.3 pp, medium confidence |
| DeepSeek V3.2 (Thinking) | 5.1% | 65.0% | estimated ± 6.3 pp, medium confidence |
| Gemini 3 Pro | 14.3% | 69.9% | estimated ± 6.3 pp, medium confidence |
| GLM-4.6 | 3.1% | 63.8% | estimated ± 6.3 pp, medium confidence |
| GLM-5 (Reasoning) | 23.4% | 73.6% | estimated ± 6.3 pp, medium confidence |
| GPT-5.1-Codex | 13.1% | 69.3% | estimated ± 6.3 pp, medium confidence |
| GPT-5.1-Codex-Max | 22.2% | 73.2% | estimated ± 6.3 pp, medium confidence |
| GPT-5.2 | 53.5% | 80.2% | estimated ± 6.3 pp, medium confidence |
| GPT-5 mini | 14.2% | 69.8% | estimated ± 6.3 pp, medium confidence |
| Grok 4.1 Fast (Reasoning) | 1.2% | 62.6% | estimated ± 6.3 pp, medium confidence |
| Grok 4 Fast (Reasoning) | 0.0% | 61.8% | estimated ± 6.3 pp, medium confidence |
| MiniMax M2.5 | 14.9% | 70.1% | estimated ± 6.3 pp, medium confidence |
| Mistral Large 4 | 78.4% | 82.1% | estimated ± 6.3 pp, low confidence |
| Qwen3.5 Plus | 15.7% | 70.5% | estimated ± 6.3 pp, medium confidence |
| Qwen3 Max | 3.5% | 64.1% | estimated ± 6.3 pp, medium confidence |