Calibration
CritPt → HealthBench Hard
HealthBench Hard is estimated from CritPt with a Michaelis–Menten curve fitted on 11 models measured on both: y = 0.3755·x / (0.02925 + x), R² = 0.40, cross-validated error 9.0 pp. It is used for 3 estimates.
| Estimated model | CritPt | HealthBench Hard | Source |
|---|---|---|---|
| Gemini 3 Pro Deep Think | 25.7% | 33.7% | estimated ± 9.0 pp, low confidence |
| GPT-5.4 Pro | 30.0% | 34.2% | estimated ± 9.0 pp, low confidence |
| GPT-5.5 Pro | 30.6% | 34.3% | estimated ± 9.0 pp, low confidence |