benchgap
Knowledge & reasoning

Humanity's Last Exam leaderboard

As of 2026-10-07, the highest measured score on Humanity's Last Exam is 61.4% by Claude Opus 5.5 (max with fallback). 6 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Claude Opus 5.5 (max with fallback)61.4%measured
2Claude Fable 5.1 (max with fallback)59.1%measured
3Claude Fable 5.1 (xhigh with fallback)58.7%measured
4Claude Opus 5.5 (xhigh with fallback)57.5%measured
5Gemini 4 Argon (high)57.1%measured
6GPT-6 Astra (xhigh)56.8%estimated ± 3.4 pp, high confidence
7GPT-6.1 Sol (xhigh)56.1%estimated ± 4.1 pp, high confidence
8Claude Fable 5.1 (high with fallback)55.9%measured
9Claude Opus 5.5 (high with fallback)55.6%measured
10Claude Fable 5 (with fallback)55.5%measured
11Claude Sonnet 5.5 (max with fallback)55.0%measured
12GPT-6 Astra (max)54.7%measured
13GPT-5.6 Sol (max)54.5%estimated ± 3.4 pp, high confidence
14GPT-6.1 Sol (max)52.9%measured
15GPT-6 Astra (high)50.4%estimated ± 4.6 pp, high confidence
16Grok 4.6 (high)50.4%estimated ± 4.6 pp, high confidence
17MiMo-V2.6-Pro49.4%measured
18Gemini 3.7 Flash (high)49.4%estimated ± 4.6 pp, high confidence
19Muse Spark 1.3 (max)48.7%measured
20GPT-6 Sol (max)47.9%measured
21Gemini 3.8 Flash (high)47.8%measured
22Kimi K3 (max)46.9%measured
23Step 5 Preview46.5%measured
24Qwen3.8 Max (0902)43.1%measured
25Grok 4.7 (xhigh)43.1%measured
26GLM-5.3 (max)42.3%measured
27GLM-5.3-Flash39.9%measured
28DeepSeek V4.1 Flash (max)39.2%measured
29MiniMax-M339.0%measured
30GPT-6 Luna (max)38.5%measured
31Mistral Large 4 Preview35.0%measured
32Qwen3.8 27B (xhigh)33.9%measured
33K2 Horizon 375B A23B32.0%measured
34Inkling (xhigh)31.9%measured
35Nemotron 3 Ultra28.4%measured
36Muse Glimmer (high)22.0%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents