benchgap
Calibration

CritPt → HealthBench Hard

HealthBench Hard is estimated from CritPt with a Michaelis–Menten curve fitted on 11 models measured on both: y = 0.3755·x / (0.02925 + x), R² = 0.40, cross-validated error 9.0 pp. It is used for 3 estimates.

Estimated modelCritPtHealthBench HardSource
Gemini 3 Pro Deep Think25.7%33.7%estimated ± 9.0 pp, low confidence
GPT-5.4 Pro30.0%34.2%estimated ± 9.0 pp, low confidence
GPT-5.5 Pro30.6%34.3%estimated ± 9.0 pp, low confidence