benchgap
Calibration

Vals LiveCodeBench → LiveCodeBench v6

LiveCodeBench v6 is estimated from Vals LiveCodeBench with a Michaelis–Menten + offset curve fitted on 6 models measured on both: y = 0.0000 + 2.0000·x / (1.06966 + x), R² = 0.41, cross-validated error 2.4 pp. It is used for 43 estimates.

Estimated modelVals LiveCodeBenchLiveCodeBench v6Source
Claude Fable 5.190.5%91.7%estimated ± 2.4 pp, low confidence
Claude Fable 589.8%91.3%estimated ± 2.4 pp, low confidence
Gemini 3.8 Flash89.5%91.1%estimated ± 2.4 pp, low confidence
Claude Opus 589.0%90.8%estimated ± 2.4 pp, low confidence
Gemini 3.7 Flash88.7%90.7%estimated ± 2.4 pp, low confidence
Gemini 3.1 Pro88.5%90.6%estimated ± 2.4 pp, low confidence
Grok 4.688.2%90.4%estimated ± 2.4 pp, low confidence
Gemini 3.6 Flash88.1%90.3%estimated ± 2.4 pp, low confidence
GPT-5.2-Codex88.0%90.3%estimated ± 2.4 pp, low confidence
Claude Opus 4.887.8%90.2%estimated ± 2.4 pp, low confidence
Gemini 3.5 Flash87.6%90.0%estimated ± 2.4 pp, low confidence
Grok 4.587.4%89.9%estimated ± 2.4 pp, low confidence
GPT-5.3 Codex87.3%89.9%estimated ± 2.4 pp, low confidence
Kimi K387.2%89.8%estimated ± 2.4 pp, low confidence
GPT-5.6 Terra85.9%89.1%estimated ± 2.4 pp, low confidence
Inkling-Small85.9%89.1%estimated ± 2.4 pp, low confidence
Muse Spark 1.185.9%89.1%estimated ± 2.4 pp, low confidence
Gemini 3 Flash85.6%88.9%estimated ± 2.4 pp, low confidence
GPT-5.1-Codex85.6%88.9%estimated ± 2.4 pp, low confidence
Inkling85.5%88.8%estimated ± 2.4 pp, low confidence
GPT-5.585.3%88.7%estimated ± 2.4 pp, low confidence
Claude Opus 4.785.1%88.6%estimated ± 2.4 pp, low confidence
Grok 4.384.5%88.3%estimated ± 2.4 pp, low confidence
Grok 4.2084.3%88.1%estimated ± 2.4 pp, low confidence
GPT-5.4 nano84.0%88.0%estimated ± 2.4 pp, low confidence
Ling 3.0 Flash84.0%88.0%estimated ± 2.4 pp, low confidence
GPT-5.1-Codex-Max83.6%87.7%estimated ± 2.4 pp, low confidence
Qwen3.5 Flash83.3%87.6%estimated ± 2.4 pp, low confidence
GPT-5.6 Sol82.6%87.1%estimated ± 2.4 pp, low confidence
Claude Sonnet 582.4%87.0%estimated ± 2.4 pp, low confidence
Claude Sonnet 4.682.1%86.8%estimated ± 2.4 pp, low confidence
Kimi K2.7 Code82.1%86.8%estimated ± 2.4 pp, low confidence
GPT-5.4 mini81.5%86.5%estimated ± 2.4 pp, low confidence
MiMo-V2.581.5%86.5%estimated ± 2.4 pp, low confidence
MiMo-V2.5-Pro81.4%86.4%estimated ± 2.4 pp, low confidence
GLM-4.681.0%86.2%estimated ± 2.4 pp, low confidence
Gemini 3.1 Flash-Lite80.1%85.6%estimated ± 2.4 pp, low confidence
Gemini 3.5 Flash-Lite79.0%85.0%estimated ± 2.4 pp, low confidence
Laguna M.168.1%77.8%estimated ± 2.4 pp, low confidence
Laguna XS.267.8%77.6%estimated ± 2.4 pp, low confidence
GLM-4.567.4%77.3%estimated ± 2.4 pp, low confidence
Grok Code Fast 162.0%73.4%estimated ± 2.4 pp, low confidence
Claude Haiku 4.541.2%55.6%estimated ± 2.4 pp, low confidence