benchgap
Math

AIME 2024 leaderboard

As of 2026-10-10, the highest measured score on AIME 2024 is 95.8% by GPT-OSS 120B. 28 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-5.297.4%estimated ± 2.2 pp, low confidence
2Step 3.5 Flash96.3%estimated ± 2.2 pp, low confidence
3MAI-Thinking-196.1%estimated ± 2.2 pp, low confidence
4GPT-OSS 120B95.8%measured
5Kimi K2.595.7%estimated ± 2.2 pp, low confidence
6Kimi K2.5 (Reasoning)95.7%estimated ± 2.2 pp, low confidence
7GLM-4.795.6%estimated ± 2.2 pp, low confidence
8Ternary Bonsai 2 27B95.3%estimated ± 2.2 pp, low confidence
9MiMo-V2-Flash94.9%estimated ± 2.2 pp, low confidence
10DeepSeek V3.2 (Thinking)94.4%estimated ± 2.2 pp, low confidence
11o4-mini (high)94.2%estimated ± 2.2 pp, low confidence
12Qwen3 235B 2507 (Reasoning)94.0%estimated ± 2.2 pp, medium confidence
13GLM-4.7-Flash93.7%estimated ± 2.2 pp, medium confidence
14DeepSeek V3.1 (Reasoning)93.1%measured
15Nemotron 3 Super 100B93.1%estimated ± 2.2 pp, medium confidence
16Granite 4.2 30B92.6%estimated ± 2.2 pp, medium confidence
17Nemotron 3 Nano 30B92.6%estimated ± 2.2 pp, medium confidence
18Sarvam 105B92.2%estimated ± 2.2 pp, medium confidence
19GPT-OSS 20B92.1%measured
20Gemini 2.5 Pro92.0%measured
21Claude Sonnet 4.591.5%estimated ± 2.2 pp, medium confidence
22Granite 4.2 8B91.4%estimated ± 2.2 pp, medium confidence
23MiniCPM5-2B91.3%estimated ± 2.2 pp, medium confidence
24Exaone 4.0 32B90.7%estimated ± 2.2 pp, medium confidence
25Ministral 3 14B (Reasoning)89.8%measured
26Grok 3 Mini89.5%measured
27Nemotron 3 Nano Omni 30B A3B89.1%estimated ± 2.2 pp, medium confidence
28o3-mini87.3%measured
29Granite 4.2 3B87.1%estimated ± 2.2 pp, medium confidence
30MiniMax M1 80k86.0%measured
31Nemotron Ultra 253B83.8%estimated ± 2.2 pp, medium confidence
32Qwen3 235B 250782.5%estimated ± 2.2 pp, medium confidence
33DeepSeek-R179.8%measured
34LFM2.5-2.6B69.7%estimated ± 2.2 pp, medium confidence
35Kimi K269.6%measured
36DeepSeek V3.166.3%measured
37LFM2.5-8B-A1B61.6%estimated ± 2.2 pp, low confidence
38MiniCPM5-1B59.7%estimated ± 2.2 pp, low confidence
39Grok 3 [Beta]52.2%measured
40Claude 4 Sonnet52.1%estimated ± 2.2 pp, low confidence
41DeepSeek V339.2%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General