benchgap
Math

HMMT Nov 2025 leaderboard

As of 2026-10-07, the highest measured score on HMMT Nov 2025 is 96.9% by GLM-5. 43 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude 3.5 Sonnet99.9%estimated ± 2.6 pp, low confidence
2GLM-4.799.9%estimated ± 2.6 pp, low confidence
3GPT-4.199.9%estimated ± 2.6 pp, low confidence
4Grok 3 [Beta]99.9%estimated ± 2.6 pp, low confidence
5Kimi K299.9%estimated ± 2.6 pp, low confidence
6Qwen3 235B 2507 (Reasoning)99.9%estimated ± 2.6 pp, low confidence
7Qwen3.5 Flash99.9%estimated ± 2.6 pp, low confidence
8GLM-596.9%measured
9GPT-5.4 mini96.8%estimated ± 2.6 pp, low confidence
10Claude Haiku 4.596.8%estimated ± 2.6 pp, low confidence
11Grok 496.8%estimated ± 2.6 pp, low confidence
12o396.8%estimated ± 2.6 pp, low confidence
13Qwen3.5 Plus96.8%estimated ± 2.6 pp, low confidence
14DeepSeek V3.296.8%estimated ± 2.6 pp, medium confidence
15GLM-4.696.7%estimated ± 2.6 pp, medium confidence
16Qwen3.6 Plus94.6%measured
17Kimi K2.694.4%estimated ± 2.0 pp, high confidence
18GLM-5.294.4%measured
19GLM-5.194.0%measured
20Claude Sonnet 4.593.3%estimated ± 2.6 pp, medium confidence
21Gemini 2.5 Flash93.3%estimated ± 2.6 pp, medium confidence
22Gemini 2.5 Pro93.3%estimated ± 2.6 pp, medium confidence
23Gemini 3 Flash93.3%estimated ± 2.6 pp, medium confidence
24Qwen 3.6 Max (preview)93.3%estimated ± 2.6 pp, medium confidence
25Claude Opus 4.593.3%measured
26GPT-5.4 nano93.2%estimated ± 2.6 pp, medium confidence
27o4-mini (high)93.2%estimated ± 2.6 pp, medium confidence
28Claude Opus 4.693.2%estimated ± 2.6 pp, low confidence
29Claude Opus 4.793.2%estimated ± 2.6 pp, low confidence
30Claude Opus 4.893.2%estimated ± 2.6 pp, low confidence
31Claude Sonnet 4.693.2%estimated ± 2.6 pp, medium confidence
32Gemini 3.1 Pro93.2%estimated ± 2.6 pp, low confidence
33Gemini 3.5 Flash93.2%estimated ± 2.6 pp, low confidence
34Gemini 3 Pro93.2%estimated ± 2.6 pp, low confidence
35GPT-5.193.2%estimated ± 2.6 pp, medium confidence
36GPT-5.293.2%estimated ± 2.6 pp, low confidence
37GPT-5.493.2%estimated ± 2.6 pp, low confidence
38GPT-5.4 Pro93.2%estimated ± 2.6 pp, low confidence
39GPT-5.593.2%estimated ± 2.6 pp, low confidence
40GPT-5.5 Pro93.2%estimated ± 2.6 pp, low confidence
41GPT-5.6 Luna93.2%estimated ± 2.6 pp, low confidence
42GPT-5.6 Sol93.2%estimated ± 2.6 pp, low confidence
43GPT-5.6 Terra93.2%estimated ± 2.6 pp, low confidence
44GPT-6 Astra93.2%estimated ± 2.6 pp, low confidence
45Muse Spark93.2%estimated ± 2.6 pp, low confidence
46Qwen3.5 397B92.7%measured
47Kimi K2.591.1%measured
48Qwen3.6-27B90.7%measured
49Qwen3.6-35B-A3B89.1%measured
50Granite 4.2 30B87.7%estimated ± 1.9 pp, low confidence
51Granite 4.2 8B77.8%estimated ± 1.9 pp, low confidence
52Granite 4.2 3B67.1%estimated ± 1.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General