Calibration
Vals LiveCodeBench → LiveCodeBench v6
LiveCodeBench v6 is estimated from Vals LiveCodeBench with a Michaelis–Menten + offset curve fitted on 6 models measured on both: y = 0.0000 + 2.0000·x / (1.06966 + x), R² = 0.41, cross-validated error 2.4 pp. It is used for 43 estimates.
| Estimated model | Vals LiveCodeBench | LiveCodeBench v6 | Source |
|---|---|---|---|
| Claude Fable 5.1 | 90.5% | 91.7% | estimated ± 2.4 pp, low confidence |
| Claude Fable 5 | 89.8% | 91.3% | estimated ± 2.4 pp, low confidence |
| Gemini 3.8 Flash | 89.5% | 91.1% | estimated ± 2.4 pp, low confidence |
| Claude Opus 5 | 89.0% | 90.8% | estimated ± 2.4 pp, low confidence |
| Gemini 3.7 Flash | 88.7% | 90.7% | estimated ± 2.4 pp, low confidence |
| Gemini 3.1 Pro | 88.5% | 90.6% | estimated ± 2.4 pp, low confidence |
| Grok 4.6 | 88.2% | 90.4% | estimated ± 2.4 pp, low confidence |
| Gemini 3.6 Flash | 88.1% | 90.3% | estimated ± 2.4 pp, low confidence |
| GPT-5.2-Codex | 88.0% | 90.3% | estimated ± 2.4 pp, low confidence |
| Claude Opus 4.8 | 87.8% | 90.2% | estimated ± 2.4 pp, low confidence |
| Gemini 3.5 Flash | 87.6% | 90.0% | estimated ± 2.4 pp, low confidence |
| Grok 4.5 | 87.4% | 89.9% | estimated ± 2.4 pp, low confidence |
| GPT-5.3 Codex | 87.3% | 89.9% | estimated ± 2.4 pp, low confidence |
| Kimi K3 | 87.2% | 89.8% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Terra | 85.9% | 89.1% | estimated ± 2.4 pp, low confidence |
| Inkling-Small | 85.9% | 89.1% | estimated ± 2.4 pp, low confidence |
| Muse Spark 1.1 | 85.9% | 89.1% | estimated ± 2.4 pp, low confidence |
| Gemini 3 Flash | 85.6% | 88.9% | estimated ± 2.4 pp, low confidence |
| GPT-5.1-Codex | 85.6% | 88.9% | estimated ± 2.4 pp, low confidence |
| Inkling | 85.5% | 88.8% | estimated ± 2.4 pp, low confidence |
| GPT-5.5 | 85.3% | 88.7% | estimated ± 2.4 pp, low confidence |
| Claude Opus 4.7 | 85.1% | 88.6% | estimated ± 2.4 pp, low confidence |
| Grok 4.3 | 84.5% | 88.3% | estimated ± 2.4 pp, low confidence |
| Grok 4.20 | 84.3% | 88.1% | estimated ± 2.4 pp, low confidence |
| GPT-5.4 nano | 84.0% | 88.0% | estimated ± 2.4 pp, low confidence |
| Ling 3.0 Flash | 84.0% | 88.0% | estimated ± 2.4 pp, low confidence |
| GPT-5.1-Codex-Max | 83.6% | 87.7% | estimated ± 2.4 pp, low confidence |
| Qwen3.5 Flash | 83.3% | 87.6% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Sol | 82.6% | 87.1% | estimated ± 2.4 pp, low confidence |
| Claude Sonnet 5 | 82.4% | 87.0% | estimated ± 2.4 pp, low confidence |
| Claude Sonnet 4.6 | 82.1% | 86.8% | estimated ± 2.4 pp, low confidence |
| Kimi K2.7 Code | 82.1% | 86.8% | estimated ± 2.4 pp, low confidence |
| GPT-5.4 mini | 81.5% | 86.5% | estimated ± 2.4 pp, low confidence |
| MiMo-V2.5 | 81.5% | 86.5% | estimated ± 2.4 pp, low confidence |
| MiMo-V2.5-Pro | 81.4% | 86.4% | estimated ± 2.4 pp, low confidence |
| GLM-4.6 | 81.0% | 86.2% | estimated ± 2.4 pp, low confidence |
| Gemini 3.1 Flash-Lite | 80.1% | 85.6% | estimated ± 2.4 pp, low confidence |
| Gemini 3.5 Flash-Lite | 79.0% | 85.0% | estimated ± 2.4 pp, low confidence |
| Laguna M.1 | 68.1% | 77.8% | estimated ± 2.4 pp, low confidence |
| Laguna XS.2 | 67.8% | 77.6% | estimated ± 2.4 pp, low confidence |
| GLM-4.5 | 67.4% | 77.3% | estimated ± 2.4 pp, low confidence |
| Grok Code Fast 1 | 62.0% | 73.4% | estimated ± 2.4 pp, low confidence |
| Claude Haiku 4.5 | 41.2% | 55.6% | estimated ± 2.4 pp, low confidence |