benchgap
Calibration

Humanity's Last Exam → GPQA Diamond

GPQA Diamond is estimated from Humanity's Last Exam with a Michaelis–Menten curve fitted on 14 models measured on both: y = 1.0536·x / (0.05994 + x), R² = 0.90, cross-validated error 1.3 pp. It is used for 16 estimates.

Estimated modelHumanity's Last ExamGPQA DiamondSource
Claude Fable 5.1 (xhigh with fallback)58.7%95.6%estimated ± 1.3 pp, high confidence
Claude Fable 5.1 (high with fallback)55.9%95.2%estimated ± 1.3 pp, high confidence
Claude Sonnet 5.5 (max with fallback)55.0%95.0%estimated ± 1.3 pp, high confidence
Claude Opus 5.5 (max with fallback)61.4%96.0%estimated ± 1.3 pp, medium confidence
Claude Opus 5.5 (xhigh with fallback)57.5%95.4%estimated ± 1.3 pp, high confidence
Gemini 4 Argon (high)57.1%95.3%estimated ± 1.3 pp, high confidence
Claude Opus 5.5 (high with fallback)55.6%95.1%estimated ± 1.3 pp, high confidence
GPT-6.1 Sol (max)52.9%94.6%estimated ± 1.3 pp, high confidence
GPT-6 Sol (max)47.9%93.6%estimated ± 1.3 pp, high confidence
MiMo-V2.6-Pro49.4%94.0%estimated ± 1.3 pp, high confidence
Step 5 Preview46.5%93.3%estimated ± 1.3 pp, high confidence
DeepSeek V4.1 Flash (max)39.2%91.4%estimated ± 1.3 pp, high confidence
Mistral Large 4 Preview35.0%90.0%estimated ± 1.3 pp, high confidence
Grok 4.7 (xhigh)43.1%92.5%estimated ± 1.3 pp, high confidence
GPT-6 Luna (max)38.5%91.2%estimated ± 1.3 pp, high confidence
Claude Fable 5 (with fallback)55.5%95.1%estimated ± 1.3 pp, high confidence