benchgap
Calibration

PostTrainBench v1.1 → ProgramBench

ProgramBench is estimated from PostTrainBench v1.1 with a offset logistic curve fitted on 5 models measured on both: y = 0.0000 + (0.9092 − 0.0000) / (1 + exp(−200.00·(x − 0.3121))), R² = 0.94, cross-validated error 4.8 pp. It is used for 9 estimates.

Estimated modelPostTrainBench v1.1ProgramBenchSource
Claude Opus 4.728.6%0.5%estimated ± 4.8 pp, low confidence
Claude Opus 4.832.9%87.9%estimated ± 4.8 pp, medium confidence
Gemini 3.1 Pro22.0%0.0%estimated ± 4.8 pp, low confidence
Gemini 4 Argon45.3%90.9%estimated ± 4.8 pp, medium confidence
GPT-5.419.0%0.0%estimated ± 4.8 pp, low confidence
GPT-5.527.2%0.0%estimated ± 4.8 pp, low confidence
GPT-5.6 Sol36.2%90.9%estimated ± 4.8 pp, medium confidence
GPT-6 Astra44.3%90.9%estimated ± 4.8 pp, medium confidence
Grok 4.523.5%0.0%estimated ± 4.8 pp, low confidence