Calibration
Vals LiveCodeBench → VulcanBench v3
VulcanBench v3 is estimated from Vals LiveCodeBench with a offset logistic curve fitted on 12 models measured on both: y = 0.7613 + (1.2000 − 0.7613) / (1 + exp(−10.44·(x − 1.0015))), R² = 0.31, cross-validated error 5.6 pp. It is used for 37 estimates.
| Estimated model | Vals LiveCodeBench | VulcanBench v3 | Source |
|---|---|---|---|
| Claude Opus 4.7 | 85.1% | 83.7% | estimated ± 5.6 pp, low confidence |
| Claude Sonnet 4.6 | 82.1% | 81.9% | estimated ± 5.6 pp, low confidence |
| DeepSeek V4 Pro 0813 | 87.5% | 85.4% | estimated ± 5.6 pp, low confidence |
| Gemini 3.1 Flash-Lite | 80.1% | 80.9% | estimated ± 5.6 pp, low confidence |
| Gemini 3.1 Pro | 88.5% | 86.2% | estimated ± 5.6 pp, low confidence |
| Gemini 3.5 Flash-Lite | 79.0% | 80.5% | estimated ± 5.6 pp, low confidence |
| Gemini 3.7 Flash | 88.7% | 86.3% | estimated ± 5.6 pp, low confidence |
| Gemini 3 Flash | 85.6% | 84.0% | estimated ± 5.6 pp, low confidence |
| GLM-4.5 | 67.4% | 77.5% | estimated ± 5.6 pp, low confidence |
| GLM-4.6 | 81.0% | 81.4% | estimated ± 5.6 pp, low confidence |
| GLM-4.7 | 82.2% | 82.0% | estimated ± 5.6 pp, low confidence |
| GLM-5.1 | 81.4% | 81.6% | estimated ± 5.6 pp, low confidence |
| GLM-5.3-Flash | 80.5% | 81.1% | estimated ± 5.6 pp, low confidence |
| GPT-5.1-Codex | 85.6% | 84.0% | estimated ± 5.6 pp, low confidence |
| GPT-5.1-Codex-Max | 83.6% | 82.7% | estimated ± 5.6 pp, low confidence |
| GPT-5.2-Codex | 88.0% | 85.8% | estimated ± 5.6 pp, low confidence |
| GPT-5.3 Codex | 87.3% | 85.2% | estimated ± 5.6 pp, low confidence |
| GPT-5.4 mini | 81.5% | 81.6% | estimated ± 5.6 pp, low confidence |
| GPT-5.4 nano | 84.0% | 83.0% | estimated ± 5.6 pp, low confidence |
| Grok 4.20 | 84.3% | 83.2% | estimated ± 5.6 pp, low confidence |
| Grok 4.3 | 84.5% | 83.3% | estimated ± 5.6 pp, low confidence |
| Grok Code Fast 1 | 62.0% | 76.9% | estimated ± 5.6 pp, low confidence |
| Inkling | 85.5% | 83.9% | estimated ± 5.6 pp, low confidence |
| Inkling-Small | 85.9% | 84.2% | estimated ± 5.6 pp, low confidence |
| Kimi K2.6 | 86.8% | 84.9% | estimated ± 5.6 pp, low confidence |
| Laguna M.1 | 68.1% | 77.6% | estimated ± 5.6 pp, low confidence |
| Laguna XS.2 | 67.8% | 77.6% | estimated ± 5.6 pp, low confidence |
| Ling 3.0 Flash | 84.0% | 83.0% | estimated ± 5.6 pp, low confidence |
| MiMo-V2.5 | 81.5% | 81.6% | estimated ± 5.6 pp, low confidence |
| MiMo-V2.5-Pro | 81.4% | 81.6% | estimated ± 5.6 pp, low confidence |
| MiniMax M2.7 | 79.9% | 80.9% | estimated ± 5.6 pp, low confidence |
| MiniMax M3 | 82.2% | 82.0% | estimated ± 5.6 pp, low confidence |
| Muse Spark 1.1 | 85.9% | 84.2% | estimated ± 5.6 pp, low confidence |
| Nemotron 3 Ultra | 86.0% | 84.3% | estimated ± 5.6 pp, low confidence |
| Qwen3.5 Flash | 83.3% | 82.6% | estimated ± 5.6 pp, low confidence |
| Qwen3.6 Plus | 86.0% | 84.3% | estimated ± 5.6 pp, low confidence |
| Qwen3.7 Max | 87.1% | 85.1% | estimated ± 5.6 pp, low confidence |