benchgap
Calibration

Vals LiveCodeBench → SciCode

SciCode is estimated from Vals LiveCodeBench with a Hill curve fitted on 7 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.91637^6.00 + x^6.00), R² = 0.67, cross-validated error 2.9 pp. It is used for 51 estimates.

Estimated modelVals LiveCodeBenchSciCodeSource
Claude Fable 589.8%56.4%estimated ± 2.9 pp, low confidence
Claude Fable 5.190.5%57.8%estimated ± 2.9 pp, low confidence
Claude Haiku 4.541.2%1.0%estimated ± 2.9 pp, low confidence
Claude Opus 4.785.1%46.9%estimated ± 2.9 pp, medium confidence
Claude Opus 4.887.8%52.3%estimated ± 2.9 pp, low confidence
Claude Opus 589.0%54.8%estimated ± 2.9 pp, low confidence
Claude Sonnet 4.682.1%40.9%estimated ± 2.9 pp, low confidence
Claude Sonnet 582.4%41.5%estimated ± 2.9 pp, low confidence
DeepSeek V4 Flash 073187.3%51.3%estimated ± 2.9 pp, medium confidence
DeepSeek V4 Pro 081387.5%51.7%estimated ± 2.9 pp, medium confidence
Gemini 3.1 Flash-Lite80.1%37.0%estimated ± 2.9 pp, low confidence
Gemini 3.1 Pro88.5%53.8%estimated ± 2.9 pp, low confidence
Gemini 3.5 Flash-Lite79.0%34.9%estimated ± 2.9 pp, low confidence
Gemini 3.6 Flash88.1%52.9%estimated ± 2.9 pp, low confidence
Gemini 3.7 Flash88.7%54.2%estimated ± 2.9 pp, low confidence
Gemini 3.8 Flash89.5%55.8%estimated ± 2.9 pp, low confidence
Gemini 3 Flash85.6%47.9%estimated ± 2.9 pp, medium confidence
GLM-4.567.4%16.4%estimated ± 2.9 pp, low confidence
GLM-4.681.0%38.8%estimated ± 2.9 pp, low confidence
GLM-4.782.2%41.1%estimated ± 2.9 pp, low confidence
GLM-5.181.4%39.5%estimated ± 2.9 pp, low confidence
GLM-5.269.5%19.2%estimated ± 2.9 pp, low confidence
GLM-5.380.5%37.8%estimated ± 2.9 pp, low confidence
GLM-5.3-Flash80.5%37.8%estimated ± 2.9 pp, low confidence
GPT-5.1-Codex85.6%47.9%estimated ± 2.9 pp, medium confidence
GPT-5.1-Codex-Max83.6%43.9%estimated ± 2.9 pp, low confidence
GPT-5.2-Codex88.0%52.7%estimated ± 2.9 pp, low confidence
GPT-5.3 Codex87.3%51.3%estimated ± 2.9 pp, medium confidence
GPT-5.4 mini81.5%39.7%estimated ± 2.9 pp, low confidence
GPT-5.4 nano84.0%44.7%estimated ± 2.9 pp, medium confidence
GPT-5.585.3%47.3%estimated ± 2.9 pp, medium confidence
GPT-5.6 Sol82.6%41.9%estimated ± 2.9 pp, low confidence
GPT-5.6 Terra85.9%48.5%estimated ± 2.9 pp, medium confidence
Grok 4.2084.3%45.3%estimated ± 2.9 pp, medium confidence
Grok 4.587.4%51.5%estimated ± 2.9 pp, medium confidence
Grok 4.688.2%53.1%estimated ± 2.9 pp, low confidence
Grok Code Fast 162.0%10.5%estimated ± 2.9 pp, low confidence
Inkling85.5%47.7%estimated ± 2.9 pp, medium confidence
Kimi K2.7 Code82.1%40.9%estimated ± 2.9 pp, low confidence
Kimi K387.2%51.1%estimated ± 2.9 pp, medium confidence
Laguna M.168.1%17.3%estimated ± 2.9 pp, low confidence
Laguna XS.267.8%16.9%estimated ± 2.9 pp, low confidence
MiMo-V2.581.5%39.7%estimated ± 2.9 pp, low confidence
MiMo-V2.5-Pro81.4%39.5%estimated ± 2.9 pp, low confidence
MiniMax M2.779.9%36.6%estimated ± 2.9 pp, low confidence
MiniMax M382.2%41.1%estimated ± 2.9 pp, low confidence
Muse Spark 1.185.9%48.5%estimated ± 2.9 pp, medium confidence
Qwen3.5 Flash83.3%43.3%estimated ± 2.9 pp, low confidence
Qwen3.6 Plus86.0%48.7%estimated ± 2.9 pp, medium confidence
Qwen3.8-27B84.0%44.7%estimated ± 2.9 pp, medium confidence
Qwen3.8 Max87.9%52.5%estimated ± 2.9 pp, low confidence