benchgap
Calibration

FrontierCode 1.1 Main → Vals LiveCodeBench

Vals LiveCodeBench is estimated from FrontierCode 1.1 Main with a linear curve fitted on 9 models measured on both: y = 0.2703·x + 0.7456, R² = 0.73, cross-validated error 1.9 pp. It is used for 7 estimates.

Estimated modelFrontierCode 1.1 MainVals LiveCodeBenchSource
Claude Haiku 5.546.4%87.1%estimated ± 1.9 pp, high confidence
Claude Opus 4.626.9%81.8%estimated ± 1.9 pp, high confidence
Claude Opus 5.554.4%89.3%estimated ± 1.9 pp, medium confidence
Claude Sonnet 5.546.2%87.0%estimated ± 1.9 pp, high confidence
GPT-6 Astra53.3%89.0%estimated ± 1.9 pp, high confidence
SWE-1.742.3%86.0%estimated ± 1.9 pp, high confidence
SWE-250.0%88.1%estimated ± 1.9 pp, high confidence