benchgap
Calibration

PostTrainBench v1.1 → DeepSWE

DeepSWE is estimated from PostTrainBench v1.1 with a Michaelis–Menten curve fitted on 8 models measured on both: y = 1.1725·x / (0.25589 + x), R² = 0.83, cross-validated error 4.4 pp. It is used for 3 estimates.

Estimated modelPostTrainBench v1.1DeepSWESource
Gemini 3.1 Pro22.0%54.2%estimated ± 4.4 pp, medium confidence
GLM-5.231.7%64.9%estimated ± 4.4 pp, high confidence
GPT-5.419.0%50.0%estimated ± 4.4 pp, medium confidence