benchgap
Math

AIME 2025 leaderboard

As of 2026-10-07, the highest measured score on AIME 2025 is 97.0% by MAI-Thinking-1. 23 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GLM-5.299.2%estimated ± 1.5 pp, low confidence
2Beam98.1%estimated ± 1.5 pp, low confidence
3A.X K297.5%estimated ± 1.5 pp, low confidence
4Inkling97.5%estimated ± 1.5 pp, low confidence
5MAI-Thinking-197.0%measured
6Kimi K2.696.9%estimated ± 1.5 pp, low confidence
7GLM-596.4%estimated ± 1.5 pp, medium confidence
8Solar Open 296.3%estimated ± 1.5 pp, medium confidence
9Inkling-Small96.1%estimated ± 1.5 pp, medium confidence
10Kimi K2.596.1%measured
11Kimi K2.5 (Reasoning)96.1%measured
12GLM-5.195.9%estimated ± 1.5 pp, medium confidence
13Qwen3.6 Plus95.9%estimated ± 1.5 pp, medium confidence
14Solar Pro 495.9%estimated ± 1.5 pp, medium confidence
15Claude Opus 4.595.8%estimated ± 1.5 pp, medium confidence
16GLM-4.795.7%measured
17Muse Glimmer 30B95.4%estimated ± 1.5 pp, medium confidence
18Ternary Bonsai 2 27B95.0%measured
19Qwen3.6-27B94.8%estimated ± 1.5 pp, medium confidence
20MiMo-V2-Flash94.1%measured
21Qwen3.5 397B94.1%estimated ± 1.5 pp, medium confidence
22Ling 3.0 Flash94.0%estimated ± 1.5 pp, medium confidence
23Qwen3.6-35B-A3B93.5%estimated ± 1.5 pp, medium confidence
24K-EXAONE 2.093.1%estimated ± 1.5 pp, medium confidence
25ZAYA1-8B89.6%estimated ± 1.5 pp, medium confidence
26Granite 4.2 30B89.2%measured
27Claude Sonnet 4.587.0%measured
28Granite 4.2 8B86.7%measured
29MiniCPM5-2B86.5%measured
30Exaone 4.0 32B85.3%measured
31Nemotron 3 Nano Omni 30B A3B82.1%measured
32Granite 4.2 3B78.3%measured
33Gemma 4 12B74.2%estimated ± 1.5 pp, medium confidence
34ZAYA1-74B-Preview72.6%estimated ± 1.5 pp, medium confidence
35LongCat-Flash-Lite-Sparse57.3%estimated ± 1.5 pp, medium confidence
36LFM2.5-2.6B51.9%measured
37LFM2.5-8B-A1B42.5%measured
38MiniCPM5-1B40.4%measured
39LLaDA2.2-mini39.1%estimated ± 1.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General