benchgap
Math

MMAnswerBench leaderboard

As of 2026-10-07, the highest measured score on MMAnswerBench is 91.0% by GLM-5.2. 72 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-5.6 Sol100.0%estimated ± 1.3 pp, low confidence
2GPT-6 Astra100.0%estimated ± 1.3 pp, low confidence
3GPT-5.6 Terra97.4%estimated ± 1.3 pp, low confidence
4GPT-5.6 Luna95.1%estimated ± 1.3 pp, low confidence
5GLM-5.291.0%measured
6GPT-5.5 Pro90.9%estimated ± 1.3 pp, low confidence
7GPT-5.4 Pro90.4%estimated ± 1.3 pp, low confidence
8GPT-5.589.9%estimated ± 1.3 pp, low confidence
9Claude Opus 4.889.0%estimated ± 1.3 pp, low confidence
10Qwen3.7 Max88.8%estimated ± 2.8 pp, low confidence
11GPT-5.488.0%estimated ± 1.3 pp, low confidence
12DeepSeek V4 Pro 081387.8%estimated ± 2.8 pp, low confidence
13Beam87.8%estimated ± 1.3 pp, high confidence
14DeepSeek V4 Flash 073187.6%estimated ± 2.8 pp, low confidence
15Claude Opus 4.787.1%estimated ± 1.3 pp, low confidence
16Claude Opus 4.687.1%estimated ± 1.3 pp, low confidence
17Qwen3.7 Plus86.6%estimated ± 2.8 pp, low confidence
18A.X K286.5%estimated ± 1.3 pp, high confidence
19Inkling86.5%estimated ± 1.3 pp, high confidence
20GPT-5.286.2%estimated ± 1.3 pp, low confidence
21Gemini 3 Pro86.2%estimated ± 1.3 pp, low confidence
22Kimi K2.686.0%measured
23Gemini 3.1 Pro85.7%estimated ± 1.3 pp, low confidence
24Muse Spark85.2%estimated ± 1.3 pp, low confidence
25Gemini 3.5 Flash85.2%estimated ± 1.3 pp, low confidence
26GPT-5.184.7%estimated ± 1.3 pp, medium confidence
27Ternary Bonsai 2 27B84.2%estimated ± 1.3 pp, high confidence
28Claude Opus 4.584.0%measured
29Solar Open 284.0%estimated ± 1.3 pp, high confidence
30GLM-5.183.8%measured
31Qwen3.6 Plus83.8%measured
32Claude Sonnet 4.683.8%estimated ± 1.3 pp, medium confidence
33Inkling-Small83.6%estimated ± 1.3 pp, high confidence
34GPT-5.4 nano83.3%estimated ± 1.3 pp, medium confidence
35o4-mini (high)83.3%estimated ± 1.3 pp, medium confidence
36Solar Pro 483.3%estimated ± 1.3 pp, high confidence
37Claude Sonnet 4.582.9%estimated ± 1.3 pp, medium confidence
38Gemini 2.5 Flash82.9%estimated ± 1.3 pp, medium confidence
39Gemini 2.5 Pro82.9%estimated ± 1.3 pp, medium confidence
40Gemini 3 Flash82.9%estimated ± 1.3 pp, medium confidence
41Qwen 3.6 Max (preview)82.9%estimated ± 1.3 pp, medium confidence
42o182.8%estimated ± 1.5 pp, low confidence
43DeepSeek V382.8%estimated ± 1.5 pp, low confidence
44GPT-4.1 mini82.8%estimated ± 1.5 pp, low confidence
45GPT-4.1 nano82.8%estimated ± 1.5 pp, low confidence
46GPT-4o82.8%estimated ± 1.5 pp, low confidence
47Llama 4 Maverick82.8%estimated ± 1.5 pp, low confidence
48Llama 4 Scout82.8%estimated ± 1.5 pp, low confidence
49GLM-582.5%measured
50GLM-4.682.4%estimated ± 1.3 pp, medium confidence
51DeepSeek V3.282.4%estimated ± 1.3 pp, medium confidence
52Claude Haiku 4.582.4%estimated ± 1.3 pp, low confidence
53Grok 482.4%estimated ± 1.3 pp, low confidence
54o382.4%estimated ± 1.3 pp, low confidence
55Qwen3.5 Plus82.4%estimated ± 1.3 pp, low confidence
56GPT-5.4 mini82.4%estimated ± 1.3 pp, low confidence
57Muse Glimmer 30B82.3%estimated ± 1.3 pp, high confidence
58Claude 3.5 Sonnet81.9%estimated ± 1.3 pp, low confidence
59GLM-4.781.9%estimated ± 1.3 pp, low confidence
60GPT-4.181.9%estimated ± 1.3 pp, low confidence
61Grok 3 [Beta]81.9%estimated ± 1.3 pp, low confidence
62Kimi K281.9%estimated ± 1.3 pp, low confidence
63Qwen3 235B 2507 (Reasoning)81.9%estimated ± 1.3 pp, low confidence
64Qwen3.5 Flash81.9%estimated ± 1.3 pp, low confidence
65MAI-Thinking-181.9%estimated ± 1.3 pp, high confidence
66Kimi K2.581.8%measured
67Qwen3.5 397B80.9%measured
68Qwen3.6-27B80.8%measured
69Ling 3.0 Flash79.7%estimated ± 1.3 pp, high confidence
70Granite 4.2 30B79.4%estimated ± 2.0 pp, low confidence
71Qwen3.6-35B-A3B78.9%measured
72K-EXAONE 2.078.2%estimated ± 1.3 pp, medium confidence
73Granite 4.2 8B73.9%estimated ± 2.0 pp, low confidence
74ZAYA1-8B73.2%estimated ± 1.3 pp, medium confidence
75MiniCPM5-2B69.3%estimated ± 1.3 pp, medium confidence
76Granite 4.2 3B67.3%estimated ± 2.0 pp, low confidence
77Gemma 4 12B57.2%estimated ± 1.3 pp, medium confidence
78ZAYA1-74B-Preview55.8%estimated ± 1.3 pp, medium confidence
79LongCat-Flash-Lite-Sparse44.0%estimated ± 1.3 pp, medium confidence
80LFM2.5-8B-A1B29.7%estimated ± 1.3 pp, medium confidence
81MiniCPM5-1B22.5%estimated ± 1.3 pp, medium confidence
82LLaDA2.2-mini18.8%estimated ± 1.3 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General