benchgap
Math

MATH-500 leaderboard

As of 2026-10-07, the highest measured score on MATH-500 is 98.8% by Ternary Bonsai 2 27B. 24 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Ternary Bonsai 2 27B98.8%measured
2GLM-5.298.2%estimated ± 3.1 pp, low confidence
3Beam98.0%estimated ± 3.1 pp, low confidence
4A.X K297.9%estimated ± 3.1 pp, low confidence
5Inkling97.9%estimated ± 3.1 pp, low confidence
6Kimi K2.697.8%estimated ± 3.1 pp, low confidence
7GLM-597.7%estimated ± 3.1 pp, medium confidence
8Kimi K2.597.7%estimated ± 3.1 pp, medium confidence
9Solar Open 297.7%estimated ± 3.1 pp, medium confidence
10Inkling-Small97.7%estimated ± 3.1 pp, medium confidence
11GLM-5.197.7%estimated ± 3.1 pp, medium confidence
12Qwen3.6 Plus97.7%estimated ± 3.1 pp, medium confidence
13Solar Pro 497.7%estimated ± 3.1 pp, medium confidence
14Claude Opus 4.597.7%estimated ± 3.1 pp, medium confidence
15Muse Glimmer 30B97.6%estimated ± 3.1 pp, medium confidence
16MAI-Thinking-197.6%estimated ± 3.1 pp, medium confidence
17Qwen3.6-27B97.5%estimated ± 3.1 pp, medium confidence
18Qwen3.5 397B97.4%estimated ± 3.1 pp, medium confidence
19Ling 3.0 Flash97.4%estimated ± 3.1 pp, medium confidence
20Qwen3.6-35B-A3B97.3%estimated ± 3.1 pp, medium confidence
21K-EXAONE 2.097.3%estimated ± 3.1 pp, medium confidence
22ZAYA1-8B96.8%estimated ± 3.1 pp, medium confidence
23LongCat-Flash-Lite-Sparse95.8%measured
24Gemma 4 12B95.2%estimated ± 3.1 pp, medium confidence
25ZAYA1-74B-Preview95.1%estimated ± 3.1 pp, medium confidence
26MiniCPM5-2B94.6%measured
27MiniCPM5-1B91.6%measured
28LLaDA2.2-mini89.5%estimated ± 3.1 pp, low confidence
29LFM2.5-8B-A1B88.8%measured
30Kanana-2 1.3B Instruct61.4%measured
31Kanana-2 3B Instruct61.2%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General