benchgap
Math

HMMT Feb 2025 leaderboard

As of 2026-10-07, the highest measured score on HMMT Feb 2025 is 97.5% by GLM-5. 26 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.7 Max100.0%estimated ± 1.8 pp, low confidence
2DeepSeek V4 Pro 081399.6%estimated ± 1.8 pp, low confidence
3DeepSeek V4 Flash 073199.4%estimated ± 1.8 pp, low confidence
4Solar Open 298.9%estimated ± 1.8 pp, low confidence
5Qwen3.7 Plus98.4%estimated ± 1.8 pp, low confidence
6Beam98.4%estimated ± 1.9 pp, low confidence
7Kimi K2.698.3%estimated ± 1.8 pp, low confidence
8A.X K297.6%estimated ± 1.9 pp, low confidence
9Inkling97.6%estimated ± 1.9 pp, low confidence
10GLM-597.5%measured
11Inkling-Small96.9%estimated ± 1.8 pp, low confidence
12Qwen3.6 Plus96.7%measured
13Ternary Bonsai 2 27B96.1%estimated ± 1.9 pp, low confidence
14GLM-5.295.5%estimated ± 1.6 pp, medium confidence
15Solar Pro 495.5%estimated ± 1.9 pp, medium confidence
16Kimi K2.595.4%measured
17GLM-5.195.3%estimated ± 1.6 pp, medium confidence
18Ling 3.0 Flash95.1%estimated ± 1.8 pp, low confidence
19Qwen3.5 397B94.8%measured
20Muse Glimmer 30B94.7%estimated ± 1.9 pp, medium confidence
21MAI-Thinking-193.9%estimated ± 1.8 pp, low confidence
22Qwen3.6-27B93.8%measured
23Claude Opus 4.592.9%measured
24Qwen3.6-35B-A3B90.7%measured
25K-EXAONE 2.089.9%estimated ± 1.8 pp, low confidence
26Granite 4.2 30B89.2%measured
27ZAYA1-8B85.5%estimated ± 1.8 pp, low confidence
28MiniCPM5-2B79.9%estimated ± 1.8 pp, low confidence
29Granite 4.2 8B78.3%measured
30Granite 4.2 3B66.7%measured
31Gemma 4 12B63.5%estimated ± 1.9 pp, low confidence
32ZAYA1-74B-Preview61.0%estimated ± 1.9 pp, low confidence
33LongCat-Flash-Lite-Sparse59.4%estimated ± 1.8 pp, low confidence
34MiniCPM5-1B42.3%estimated ± 1.8 pp, low confidence
35LFM2.5-8B-A1B9.0%estimated ± 1.9 pp, low confidence
36LLaDA2.2-mini1.1%estimated ± 1.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General