benchgap
Math

HMMT Feb 2026 leaderboard

As of 2026-10-07, the highest measured score on HMMT Feb 2026 is 97.1% by Qwen3.7 Max. 14 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1dots3-note Preview97.3%estimated ± 9.7 pp, low confidence
2Qwen3.7 Max97.1%measured
3DeepSeek V4 Pro 081395.2%measured
4DeepSeek V4 Flash 073194.8%measured
5Solar Open 293.9%measured
6Qwen3.7 Plus92.9%measured
7Kimi K2.692.7%measured
8GLM-5.292.5%measured
9Inkling-Small90.2%measured
10Beam89.8%estimated ± 4.0 pp, high confidence
11A.X K288.9%estimated ± 4.0 pp, high confidence
12Inkling88.9%estimated ± 4.0 pp, high confidence
13Qwen3.5 397B87.9%measured
14Qwen3.6 Plus87.8%measured
15Ternary Bonsai 2 27B87.2%estimated ± 4.0 pp, high confidence
16Kimi K2.587.1%measured
17Ling 3.0 Flash87.0%measured
18Solar Pro 486.5%estimated ± 4.0 pp, high confidence
19GLM-586.4%measured
20Muse Glimmer 30B85.6%estimated ± 4.0 pp, high confidence
21Claude Opus 4.585.3%measured
22MAI-Thinking-184.9%measured
23Qwen3.6-27B84.3%measured
24Qwen3.6-35B-A3B83.6%measured
25Granite 4.2 30B83.2%estimated ± 1.2 pp, low confidence
26GLM-5.182.6%measured
27K-EXAONE 2.078.4%measured
28Granite 4.2 8B77.0%estimated ± 1.2 pp, low confidence
29ZAYA1-8B71.6%measured
30Granite 4.2 3B69.5%estimated ± 1.2 pp, low confidence
31MiniCPM5-2B63.8%measured
32Gemma 4 12B57.5%estimated ± 4.0 pp, high confidence
33ZAYA1-74B-Preview55.7%estimated ± 4.0 pp, high confidence
34LongCat-Flash-Lite-Sparse40.5%measured
35LFM2.5-8B-A1B27.3%estimated ± 4.0 pp, high confidence
36MiniCPM5-1B25.8%measured
37LLaDA2.2-mini24.2%estimated ± 4.0 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General