benchgap
Calibration

CritPt → Humanity's Last Exam

Humanity's Last Exam is estimated from CritPt with a linear curve fitted on 27 models measured on both: y = 0.9277·x + 0.2667, R² = 0.87, cross-validated error 4.1 pp. It is used for 1 estimate.

Estimated modelCritPtHumanity's Last ExamSource
GPT-6.1 Sol (xhigh)31.7%56.1%estimated ± 4.1 pp, high confidence