benchgap
Math

AIME26 leaderboard

As of 2026-10-07, the highest measured score on AIME26 is 99.2% by GLM-5.2. 58 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 4.6100.0%estimated ± 0.6 pp, low confidence
2Claude Opus 4.7100.0%estimated ± 0.6 pp, low confidence
3Claude Opus 4.8100.0%estimated ± 0.6 pp, low confidence
4Gemini 3.1 Pro100.0%estimated ± 0.6 pp, low confidence
5Gemini 3 Pro100.0%estimated ± 0.6 pp, low confidence
6GPT-5.2100.0%estimated ± 0.6 pp, low confidence
7GPT-5.4100.0%estimated ± 0.6 pp, low confidence
8GPT-5.4 Pro100.0%estimated ± 0.6 pp, low confidence
9GPT-5.5100.0%estimated ± 0.6 pp, low confidence
10GPT-5.5 Pro100.0%estimated ± 0.6 pp, low confidence
11GPT-5.6 Luna100.0%estimated ± 0.6 pp, low confidence
12GPT-5.6 Sol100.0%estimated ± 0.6 pp, low confidence
13GPT-5.6 Terra100.0%estimated ± 0.6 pp, low confidence
14GPT-6 Astra100.0%estimated ± 0.6 pp, low confidence
15GLM-5.299.2%measured
16Beam97.8%measured
17Qwen3.7 Max97.3%estimated ± 1.3 pp, medium confidence
18A.X K297.1%measured
19Inkling97.1%measured
20DeepSeek V4 Pro 081396.9%estimated ± 1.3 pp, medium confidence
21DeepSeek V4 Flash 073196.8%estimated ± 1.3 pp, medium confidence
22Muse Spark96.4%estimated ± 0.6 pp, low confidence
23Qwen3.7 Plus96.4%estimated ± 1.3 pp, high confidence
24Gemini 3.5 Flash96.4%estimated ± 0.6 pp, low confidence
25Kimi K2.696.4%measured
26Ternary Bonsai 2 27B95.8%measured
27GLM-595.8%measured
28Kimi K2.595.8%measured
29Solar Open 295.7%measured
30Inkling-Small95.5%measured
31GPT-5.195.5%estimated ± 0.6 pp, medium confidence
32Claude 3.5 Sonnet95.5%estimated ± 0.6 pp, low confidence
33Claude Haiku 4.595.5%estimated ± 0.6 pp, low confidence
34Claude Sonnet 4.595.5%estimated ± 0.6 pp, medium confidence
35Claude Sonnet 4.695.5%estimated ± 0.6 pp, medium confidence
36DeepSeek V3.295.5%estimated ± 0.6 pp, medium confidence
37Gemini 2.5 Flash95.5%estimated ± 0.6 pp, medium confidence
38Gemini 2.5 Pro95.5%estimated ± 0.6 pp, medium confidence
39Gemini 3 Flash95.5%estimated ± 0.6 pp, medium confidence
40GLM-4.695.5%estimated ± 0.6 pp, medium confidence
41GLM-4.795.5%estimated ± 0.6 pp, low confidence
42GPT-4.195.5%estimated ± 0.6 pp, low confidence
43GPT-5.4 mini95.5%estimated ± 0.6 pp, low confidence
44GPT-5.4 nano95.5%estimated ± 0.6 pp, medium confidence
45Grok 3 [Beta]95.5%estimated ± 0.6 pp, low confidence
46Grok 495.5%estimated ± 0.6 pp, low confidence
47Kimi K295.5%estimated ± 0.6 pp, low confidence
48o395.5%estimated ± 0.6 pp, low confidence
49o4-mini (high)95.5%estimated ± 0.6 pp, medium confidence
50Qwen3 235B 2507 (Reasoning)95.5%estimated ± 0.6 pp, low confidence
51Qwen3.5 Flash95.5%estimated ± 0.6 pp, low confidence
52Qwen3.5 Plus95.5%estimated ± 0.6 pp, low confidence
53Qwen 3.6 Max (preview)95.5%estimated ± 0.6 pp, medium confidence
54Kimi K2.5 (Reasoning)95.4%estimated ± 4.6 pp, medium confidence
55o195.4%estimated ± 0.6 pp, low confidence
56DeepSeek V395.4%estimated ± 0.6 pp, low confidence
57GPT-4.1 mini95.4%estimated ± 0.6 pp, low confidence
58GPT-4.1 nano95.4%estimated ± 0.6 pp, low confidence
59GPT-4o95.4%estimated ± 0.6 pp, low confidence
60Llama 4 Maverick95.4%estimated ± 0.6 pp, low confidence
61Llama 4 Scout95.4%estimated ± 0.6 pp, low confidence
62GLM-5.195.3%measured
63Qwen3.6 Plus95.3%measured
64Solar Pro 495.3%measured
65Claude Opus 4.595.1%measured
66Muse Glimmer 30B94.7%measured
67MAI-Thinking-194.5%measured
68Qwen3.6-27B94.1%measured
69MiMo-V2-Flash93.6%estimated ± 4.6 pp, medium confidence
70Qwen3.5 397B93.3%measured
71Ling 3.0 Flash93.2%measured
72Qwen3.6-35B-A3B92.7%measured
73Granite 4.2 30B92.4%estimated ± 1.0 pp, low confidence
74K-EXAONE 2.092.3%measured
75ZAYA1-8B89.1%measured
76Granite 4.2 8B87.5%estimated ± 1.0 pp, low confidence
77MiniCPM5-2B86.5%measured
78Exaone 4.0 32B85.5%estimated ± 4.6 pp, medium confidence
79Nemotron 3 Nano Omni 30B A3B82.6%estimated ± 4.6 pp, medium confidence
80Granite 4.2 3B81.4%estimated ± 1.0 pp, low confidence
81Gemma 4 12B77.5%measured
82ZAYA1-74B-Preview76.4%measured
83LongCat-Flash-Lite-Sparse65.7%measured
84LFM2.5-2.6B54.7%estimated ± 4.6 pp, medium confidence
85LFM2.5-8B-A1B50.0%measured
86MiniCPM5-1B40.4%measured
87LLaDA2.2-mini35.1%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General