benchgap
Agentic · tools

Agents' Last Exam leaderboard

As of 2026-10-07, the highest measured score on Agents' Last Exam is 59.3% by GPT-6 Astra. 48 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra59.3%measured
2GPT-6.1 Sol57.5%estimated ± 6.3 pp, medium confidence
3GPT-6 Sol56.4%measured
4Qwen3.8 Max52.4%measured
5Muse Spark 1.351.4%estimated ± 6.3 pp, medium confidence
6Qwen3.8-Flash-Next51.2%measured
7Claude Fable 5.150.8%estimated ± 6.3 pp, medium confidence
8Claude Opus 5.550.8%estimated ± 6.3 pp, medium confidence
9Claude Sonnet 5.550.2%estimated ± 6.3 pp, medium confidence
10Claude 4.1 Opus48.9%estimated ± 8.7 pp, low confidence
11Claude 4 Sonnet48.9%estimated ± 8.7 pp, low confidence
12Claude Haiku 4.548.9%estimated ± 8.7 pp, low confidence
13Claude Sonnet 4.548.9%estimated ± 8.7 pp, low confidence
14Gemini 3 Flash48.9%estimated ± 8.7 pp, low confidence
15Gemini 3 Pro48.9%estimated ± 8.7 pp, low confidence
16GPT-5.1-Codex48.9%estimated ± 8.7 pp, low confidence
17GPT-5.248.9%estimated ± 8.7 pp, low confidence
18GPT-5.2-Codex48.9%estimated ± 8.7 pp, low confidence
19GPT-5.3 Codex48.9%estimated ± 8.7 pp, low confidence
20GPT-5 (high)48.9%estimated ± 8.7 pp, low confidence
21Kimi K2.548.9%estimated ± 8.7 pp, low confidence
22Qwen3.5 Plus48.9%estimated ± 8.7 pp, low confidence
23GPT-6 Luna45.7%estimated ± 6.3 pp, medium confidence
24Kimi K344.5%estimated ± 6.3 pp, medium confidence
25Gemini 3.8 Flash42.9%estimated ± 6.3 pp, medium confidence
26Qwen3.8-27B42.9%measured
27Claude Haiku 5.542.6%estimated ± 6.3 pp, medium confidence
28Grok 4.741.3%estimated ± 6.3 pp, medium confidence
29Gemini 4 Argon39.5%measured
30Mistral Large 439.0%estimated ± 6.3 pp, medium confidence
31DeepSeek V4.1 Flash31.8%measured
32MiMo-V2.6-Pro31.6%measured
33Ling 3.1 Flash30.3%estimated ± 2.1 pp, high confidence
34Fugu Cyber30.3%estimated ± 2.1 pp, high confidence
35Atria Dawn Preview30.3%estimated ± 2.1 pp, high confidence
36Gemini 3.8 Flash Cyber30.3%estimated ± 2.1 pp, high confidence
37Step 5 Preview29.5%measured
38GPT-5.6 Sol28.8%estimated ± 2.1 pp, high confidence
39Inkling28.5%estimated ± 6.3 pp, medium confidence
40GLM-5.328.5%measured
41MiMo-V2.6-Flash27.6%measured
42Claude Mythos 527.0%estimated ± 2.1 pp, high confidence
43Gemini 3.7 Flash26.3%measured
44GLM-5.3-Flash26.3%measured
45DeepSeek V4 Pro 081325.7%measured
46Gemini 3.5 Flash Cyber25.4%estimated ± 2.1 pp, high confidence
47Claude Mythos Preview25.2%estimated ± 2.1 pp, high confidence
48DeepSeek V4 Flash 073125.2%measured
49GPT-5.524.2%estimated ± 2.1 pp, high confidence
50GPT-5.6 Terra24.2%estimated ± 2.1 pp, high confidence
51GPT-5.424.0%estimated ± 2.1 pp, high confidence
52GPT-5.6 Luna24.0%estimated ± 2.1 pp, high confidence
53Claude Opus 4.524.0%estimated ± 2.1 pp, medium confidence
54Claude Opus 4.624.0%estimated ± 2.1 pp, medium confidence
55Claude Opus 4.7 (Adaptive)24.0%estimated ± 2.1 pp, medium confidence
56Claude Sonnet 4.624.0%estimated ± 2.1 pp, medium confidence
57GLM-524.0%estimated ± 2.1 pp, medium confidence
58GLM-5.124.0%estimated ± 2.1 pp, medium confidence
59Muse Spark24.0%estimated ± 2.1 pp, medium confidence
60Muse Spark 1.124.0%estimated ± 2.1 pp, medium confidence
61Muse Glimmer 30B23.0%estimated ± 6.3 pp, low confidence
62Hy4 preview22.8%measured
63MiniMax M322.6%estimated ± 6.3 pp, low confidence
64Nemotron 3 Ultra12.2%estimated ± 6.3 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General