Calibration
Vibe Code Bench → SWE-Rebench
SWE-Rebench is estimated from Vibe Code Bench with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.6226 − 0.0000)·x^6.00 / (0.19846^6.00 + x^6.00), R² = 0.54, cross-validated error 6.5 pp. It is used for 30 estimates.
| Estimated model | Vibe Code Bench | SWE-Rebench | Source |
|---|---|---|---|
| Claude Haiku 4.5 Thinking | 11.4% | 2.2% | estimated ± 6.5 pp, low confidence |
| Claude Opus 4.5 Thinking | 20.6% | 34.7% | estimated ± 6.5 pp, low confidence |
| Claude Opus 4.6 (Adaptive) | 53.5% | 62.1% | estimated ± 6.5 pp, low confidence |
| Claude Opus 4.7 | 71.0% | 62.2% | estimated ± 6.5 pp, low confidence |
| Claude Sonnet 4.5 Thinking | 22.6% | 42.8% | estimated ± 6.5 pp, low confidence |
| DeepSeek V3.2 (Thinking) | 5.1% | 0.0% | estimated ± 6.5 pp, low confidence |
| Gemini 3.1 Flash-Lite | 0.0% | 0.0% | estimated ± 6.5 pp, low confidence |
| Gemini 3.1 Pro | 32.0% | 58.9% | estimated ± 6.5 pp, low confidence |
| Gemini 3.5 Flash | 48.7% | 62.0% | estimated ± 6.5 pp, low confidence |
| Gemini 3 Flash | 20.2% | 32.8% | estimated ± 6.5 pp, low confidence |
| Gemini 3 Pro | 14.3% | 7.6% | estimated ± 6.5 pp, low confidence |
| Gemini 4 Argon | 91.9% | 62.3% | estimated ± 6.5 pp, low confidence |
| GLM-4.6 | 3.1% | 0.0% | estimated ± 6.5 pp, low confidence |
| GLM-5 (Reasoning) | 23.4% | 45.2% | estimated ± 6.5 pp, low confidence |
| GPT-5.1 | 24.6% | 48.8% | estimated ± 6.5 pp, low confidence |
| GPT-5.1-Codex | 13.1% | 4.8% | estimated ± 6.5 pp, low confidence |
| GPT-5.1-Codex-Max | 22.2% | 41.1% | estimated ± 6.5 pp, low confidence |
| GPT-5.2-Codex | 37.9% | 61.0% | estimated ± 6.5 pp, low confidence |
| GPT-5.4 | 67.4% | 62.2% | estimated ± 6.5 pp, low confidence |
| GPT-5.4 mini | 48.0% | 62.0% | estimated ± 6.5 pp, low confidence |
| GPT-5.4 nano | 26.1% | 52.2% | estimated ± 6.5 pp, low confidence |
| GPT-5.5 | 69.8% | 62.2% | estimated ± 6.5 pp, low confidence |
| GPT-5 (high) | 20.1% | 32.3% | estimated ± 6.5 pp, low confidence |
| GPT-5 mini | 14.2% | 7.3% | estimated ± 6.5 pp, low confidence |
| Grok 4.1 Fast (Reasoning) | 1.2% | 0.0% | estimated ± 6.5 pp, low confidence |
| Grok 4 Fast (Reasoning) | 0.0% | 0.0% | estimated ± 6.5 pp, low confidence |
| MiniMax M2.5 | 14.9% | 9.3% | estimated ± 6.5 pp, low confidence |
| Mistral Large 4 | 78.4% | 62.2% | estimated ± 6.5 pp, low confidence |
| Qwen3.5 Plus | 15.7% | 12.4% | estimated ± 6.5 pp, low confidence |
| Qwen3 Max | 3.5% | 0.0% | estimated ± 6.5 pp, low confidence |