benchgap
Calibration

CritPt → BioMysteryBench (human-difficult)

BioMysteryBench (human-difficult) is estimated from CritPt with a Michaelis–Menten curve fitted on 6 models measured on both: y = 0.5327·x / (0.02161 + x), R² = 0.56, cross-validated error 6.8 pp. It is used for 7 estimates.

Estimated modelCritPtBioMysteryBench (human-difficult)Source
Claude 4.1 Opus Thinking0.0%0.0%estimated ± 6.8 pp, low confidence
Claude Haiku 5.518.9%47.8%estimated ± 6.8 pp, low confidence
Gemini 3 Pro Deep Think25.7%49.1%estimated ± 6.8 pp, low confidence
GLM-5.3-Flash15.4%46.7%estimated ± 6.8 pp, low confidence
GPT-5.4 Pro30.0%49.7%estimated ± 6.8 pp, low confidence
GPT-5.5 Pro30.6%49.8%estimated ± 6.8 pp, low confidence
Hy4 preview16.9%47.2%estimated ± 6.8 pp, low confidence