Calibration
CritPt → BioMysteryBench (human-difficult)
BioMysteryBench (human-difficult) is estimated from CritPt with a Michaelis–Menten curve fitted on 6 models measured on both: y = 0.5327·x / (0.02161 + x), R² = 0.56, cross-validated error 6.8 pp. It is used for 7 estimates.
| Estimated model | CritPt | BioMysteryBench (human-difficult) | Source |
|---|---|---|---|
| Claude 4.1 Opus Thinking | 0.0% | 0.0% | estimated ± 6.8 pp, low confidence |
| Claude Haiku 5.5 | 18.9% | 47.8% | estimated ± 6.8 pp, low confidence |
| Gemini 3 Pro Deep Think | 25.7% | 49.1% | estimated ± 6.8 pp, low confidence |
| GLM-5.3-Flash | 15.4% | 46.7% | estimated ± 6.8 pp, low confidence |
| GPT-5.4 Pro | 30.0% | 49.7% | estimated ± 6.8 pp, low confidence |
| GPT-5.5 Pro | 30.6% | 49.8% | estimated ± 6.8 pp, low confidence |
| Hy4 preview | 16.9% | 47.2% | estimated ± 6.8 pp, low confidence |