benchgap
Calibration

GPQA Diamond → Humanity's Last Exam

Humanity's Last Exam is estimated from GPQA Diamond with a linear curve fitted on 14 models measured on both: y = 2.5860·x + -1.9499, R² = 0.83, cross-validated error 4.6 pp. It is used for 3 estimates.

Estimated modelGPQA DiamondHumanity's Last ExamSource
GPT-6 Astra (high)94.9%50.4%estimated ± 4.6 pp, high confidence
Grok 4.6 (high)94.9%50.4%estimated ± 4.6 pp, high confidence
Gemini 3.7 Flash (high)94.5%49.4%estimated ± 4.6 pp, high confidence