benchgap
Calibration

Vals LiveCodeBench → VulcanBench v3

VulcanBench v3 is estimated from Vals LiveCodeBench with a offset logistic curve fitted on 12 models measured on both: y = 0.7613 + (1.2000 − 0.7613) / (1 + exp(−10.44·(x − 1.0015))), R² = 0.31, cross-validated error 5.6 pp. It is used for 37 estimates.

Estimated modelVals LiveCodeBenchVulcanBench v3Source
Claude Opus 4.785.1%83.7%estimated ± 5.6 pp, low confidence
Claude Sonnet 4.682.1%81.9%estimated ± 5.6 pp, low confidence
DeepSeek V4 Pro 081387.5%85.4%estimated ± 5.6 pp, low confidence
Gemini 3.1 Flash-Lite80.1%80.9%estimated ± 5.6 pp, low confidence
Gemini 3.1 Pro88.5%86.2%estimated ± 5.6 pp, low confidence
Gemini 3.5 Flash-Lite79.0%80.5%estimated ± 5.6 pp, low confidence
Gemini 3.7 Flash88.7%86.3%estimated ± 5.6 pp, low confidence
Gemini 3 Flash85.6%84.0%estimated ± 5.6 pp, low confidence
GLM-4.567.4%77.5%estimated ± 5.6 pp, low confidence
GLM-4.681.0%81.4%estimated ± 5.6 pp, low confidence
GLM-4.782.2%82.0%estimated ± 5.6 pp, low confidence
GLM-5.181.4%81.6%estimated ± 5.6 pp, low confidence
GLM-5.3-Flash80.5%81.1%estimated ± 5.6 pp, low confidence
GPT-5.1-Codex85.6%84.0%estimated ± 5.6 pp, low confidence
GPT-5.1-Codex-Max83.6%82.7%estimated ± 5.6 pp, low confidence
GPT-5.2-Codex88.0%85.8%estimated ± 5.6 pp, low confidence
GPT-5.3 Codex87.3%85.2%estimated ± 5.6 pp, low confidence
GPT-5.4 mini81.5%81.6%estimated ± 5.6 pp, low confidence
GPT-5.4 nano84.0%83.0%estimated ± 5.6 pp, low confidence
Grok 4.2084.3%83.2%estimated ± 5.6 pp, low confidence
Grok 4.384.5%83.3%estimated ± 5.6 pp, low confidence
Grok Code Fast 162.0%76.9%estimated ± 5.6 pp, low confidence
Inkling85.5%83.9%estimated ± 5.6 pp, low confidence
Inkling-Small85.9%84.2%estimated ± 5.6 pp, low confidence
Kimi K2.686.8%84.9%estimated ± 5.6 pp, low confidence
Laguna M.168.1%77.6%estimated ± 5.6 pp, low confidence
Laguna XS.267.8%77.6%estimated ± 5.6 pp, low confidence
Ling 3.0 Flash84.0%83.0%estimated ± 5.6 pp, low confidence
MiMo-V2.581.5%81.6%estimated ± 5.6 pp, low confidence
MiMo-V2.5-Pro81.4%81.6%estimated ± 5.6 pp, low confidence
MiniMax M2.779.9%80.9%estimated ± 5.6 pp, low confidence
MiniMax M382.2%82.0%estimated ± 5.6 pp, low confidence
Muse Spark 1.185.9%84.2%estimated ± 5.6 pp, low confidence
Nemotron 3 Ultra86.0%84.3%estimated ± 5.6 pp, low confidence
Qwen3.5 Flash83.3%82.6%estimated ± 5.6 pp, low confidence
Qwen3.6 Plus86.0%84.3%estimated ± 5.6 pp, low confidence
Qwen3.7 Max87.1%85.1%estimated ± 5.6 pp, low confidence