Calibration
Vals LiveCodeBench → SciCode
SciCode is estimated from Vals LiveCodeBench with a Hill curve fitted on 7 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.91637^6.00 + x^6.00), R² = 0.67, cross-validated error 2.9 pp. It is used for 51 estimates.
| Estimated model | Vals LiveCodeBench | SciCode | Source |
|---|---|---|---|
| Claude Fable 5 | 89.8% | 56.4% | estimated ± 2.9 pp, low confidence |
| Claude Fable 5.1 | 90.5% | 57.8% | estimated ± 2.9 pp, low confidence |
| Claude Haiku 4.5 | 41.2% | 1.0% | estimated ± 2.9 pp, low confidence |
| Claude Opus 4.7 | 85.1% | 46.9% | estimated ± 2.9 pp, medium confidence |
| Claude Opus 4.8 | 87.8% | 52.3% | estimated ± 2.9 pp, low confidence |
| Claude Opus 5 | 89.0% | 54.8% | estimated ± 2.9 pp, low confidence |
| Claude Sonnet 4.6 | 82.1% | 40.9% | estimated ± 2.9 pp, low confidence |
| Claude Sonnet 5 | 82.4% | 41.5% | estimated ± 2.9 pp, low confidence |
| DeepSeek V4 Flash 0731 | 87.3% | 51.3% | estimated ± 2.9 pp, medium confidence |
| DeepSeek V4 Pro 0813 | 87.5% | 51.7% | estimated ± 2.9 pp, medium confidence |
| Gemini 3.1 Flash-Lite | 80.1% | 37.0% | estimated ± 2.9 pp, low confidence |
| Gemini 3.1 Pro | 88.5% | 53.8% | estimated ± 2.9 pp, low confidence |
| Gemini 3.5 Flash-Lite | 79.0% | 34.9% | estimated ± 2.9 pp, low confidence |
| Gemini 3.6 Flash | 88.1% | 52.9% | estimated ± 2.9 pp, low confidence |
| Gemini 3.7 Flash | 88.7% | 54.2% | estimated ± 2.9 pp, low confidence |
| Gemini 3.8 Flash | 89.5% | 55.8% | estimated ± 2.9 pp, low confidence |
| Gemini 3 Flash | 85.6% | 47.9% | estimated ± 2.9 pp, medium confidence |
| GLM-4.5 | 67.4% | 16.4% | estimated ± 2.9 pp, low confidence |
| GLM-4.6 | 81.0% | 38.8% | estimated ± 2.9 pp, low confidence |
| GLM-4.7 | 82.2% | 41.1% | estimated ± 2.9 pp, low confidence |
| GLM-5.1 | 81.4% | 39.5% | estimated ± 2.9 pp, low confidence |
| GLM-5.2 | 69.5% | 19.2% | estimated ± 2.9 pp, low confidence |
| GLM-5.3 | 80.5% | 37.8% | estimated ± 2.9 pp, low confidence |
| GLM-5.3-Flash | 80.5% | 37.8% | estimated ± 2.9 pp, low confidence |
| GPT-5.1-Codex | 85.6% | 47.9% | estimated ± 2.9 pp, medium confidence |
| GPT-5.1-Codex-Max | 83.6% | 43.9% | estimated ± 2.9 pp, low confidence |
| GPT-5.2-Codex | 88.0% | 52.7% | estimated ± 2.9 pp, low confidence |
| GPT-5.3 Codex | 87.3% | 51.3% | estimated ± 2.9 pp, medium confidence |
| GPT-5.4 mini | 81.5% | 39.7% | estimated ± 2.9 pp, low confidence |
| GPT-5.4 nano | 84.0% | 44.7% | estimated ± 2.9 pp, medium confidence |
| GPT-5.5 | 85.3% | 47.3% | estimated ± 2.9 pp, medium confidence |
| GPT-5.6 Sol | 82.6% | 41.9% | estimated ± 2.9 pp, low confidence |
| GPT-5.6 Terra | 85.9% | 48.5% | estimated ± 2.9 pp, medium confidence |
| Grok 4.20 | 84.3% | 45.3% | estimated ± 2.9 pp, medium confidence |
| Grok 4.5 | 87.4% | 51.5% | estimated ± 2.9 pp, medium confidence |
| Grok 4.6 | 88.2% | 53.1% | estimated ± 2.9 pp, low confidence |
| Grok Code Fast 1 | 62.0% | 10.5% | estimated ± 2.9 pp, low confidence |
| Inkling | 85.5% | 47.7% | estimated ± 2.9 pp, medium confidence |
| Kimi K2.7 Code | 82.1% | 40.9% | estimated ± 2.9 pp, low confidence |
| Kimi K3 | 87.2% | 51.1% | estimated ± 2.9 pp, medium confidence |
| Laguna M.1 | 68.1% | 17.3% | estimated ± 2.9 pp, low confidence |
| Laguna XS.2 | 67.8% | 16.9% | estimated ± 2.9 pp, low confidence |
| MiMo-V2.5 | 81.5% | 39.7% | estimated ± 2.9 pp, low confidence |
| MiMo-V2.5-Pro | 81.4% | 39.5% | estimated ± 2.9 pp, low confidence |
| MiniMax M2.7 | 79.9% | 36.6% | estimated ± 2.9 pp, low confidence |
| MiniMax M3 | 82.2% | 41.1% | estimated ± 2.9 pp, low confidence |
| Muse Spark 1.1 | 85.9% | 48.5% | estimated ± 2.9 pp, medium confidence |
| Qwen3.5 Flash | 83.3% | 43.3% | estimated ± 2.9 pp, low confidence |
| Qwen3.6 Plus | 86.0% | 48.7% | estimated ± 2.9 pp, medium confidence |
| Qwen3.8-27B | 84.0% | 44.7% | estimated ± 2.9 pp, medium confidence |
| Qwen3.8 Max | 87.9% | 52.5% | estimated ± 2.9 pp, low confidence |