benchgap
Calibration

Humanity's Last Exam → CritPt

CritPt is estimated from Humanity's Last Exam with a Hill curve fitted on 27 models measured on both: y = 0.0041 + (0.3628 − 0.0041)·x^6.00 / (0.42882^6.00 + x^6.00), R² = 0.90, cross-validated error 3.5 pp. It is used for 3 estimates.

Estimated modelHumanity's Last ExamCritPtSource
Claude Fable 5.1 (high with fallback)55.9%30.2%estimated ± 3.5 pp, high confidence
Claude Opus 5.5 (high with fallback)55.6%30.0%estimated ± 3.5 pp, high confidence
Claude Fable 5 (with fallback)55.5%30.0%estimated ± 3.5 pp, high confidence