benchgap
Vision & documents

MathVision leaderboard

As of 2026-10-07, the highest measured score on MathVision is 95.2% by Qwen3.8 Max. 107 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Mythos 596.3%estimated ± 2.1 pp, low confidence
2GPT-6 Astra96.2%estimated ± 2.4 pp, low confidence
3GPT-5.6 Sol95.6%estimated ± 3.6 pp, medium confidence
4Qwen3.8 Max95.2%measured
5Kimi K394.3%measured
6Gemini 3.8 Flash93.5%estimated ± 2.1 pp, medium confidence
7Muse Spark 1.193.3%estimated ± 0.5 pp, medium confidence
8Seed 2.1 Pro92.6%measured
9GPT-5.592.1%estimated ± 3.6 pp, high confidence
10Qwen3.8-Omni-Flash91.8%measured
11Gemini 3.7 Flash91.5%estimated ± 2.1 pp, medium confidence
12GPT-5.6 Terra91.2%estimated ± 3.6 pp, high confidence
13Qwen3.8-Flash-Next90.6%measured
14Gemini 3.1 Pro90.6%estimated ± 1.4 pp, medium confidence
15Muse Glimmer 30B90.4%estimated ± 2.4 pp, medium confidence
16Kimi K2.590.3%estimated ± 3.2 pp, medium confidence
17Qwen3.7 Plus90.3%measured
18Sakana Fugu-Ultra90.1%estimated ± 2.6 pp, high confidence
19Seed 2.1 Turbo90.1%measured
20Qwen3.8-27B90.0%measured
21Claude Opus 589.7%estimated ± 4.9 pp, medium confidence
22Claude Opus 5.589.7%estimated ± 4.9 pp, medium confidence
23Gemini 3.6 Flash89.7%estimated ± 4.9 pp, medium confidence
24GPT-6.1 Sol89.7%estimated ± 4.9 pp, medium confidence
25GPT-6 Sol89.7%estimated ± 4.9 pp, medium confidence
26Qwen3.8 Max Preview89.7%estimated ± 4.9 pp, medium confidence
27Grok 4.589.7%estimated ± 4.9 pp, high confidence
28GPT-6 Luna89.7%estimated ± 4.9 pp, high confidence
29Apodex 1.189.7%estimated ± 4.9 pp, high confidence
30Apodex 1.1 Mini89.7%estimated ± 4.9 pp, high confidence
31Gemini 3.5 Flash-Lite89.7%estimated ± 4.9 pp, high confidence
32Ling 3.0 Flash VL89.7%estimated ± 4.9 pp, high confidence
33Gemini 3 Flash89.6%estimated ± 4.9 pp, high confidence
34GPT-5.3 Codex89.6%estimated ± 4.9 pp, high confidence
35DeepSeek V4.1 Flash89.4%estimated ± 4.9 pp, high confidence
36Sakana Fugu89.3%estimated ± 2.6 pp, high confidence
37GPT-5.489.2%estimated ± 1.4 pp, low confidence
38Claude Opus 4.789.1%estimated ± 4.9 pp, high confidence
39Mistral Large 489.1%estimated ± 4.9 pp, high confidence
40GPT-5.2-Codex89.0%estimated ± 4.9 pp, high confidence
41Muse Spark88.9%estimated ± 1.4 pp, low confidence
42Gemini 3.5 Flash88.7%estimated ± 2.6 pp, high confidence
43Qwen3.5 397B88.6%measured
44Holo2-235B-A22B88.4%estimated ± 2.4 pp, medium confidence
45Claude Opus 4.7 (Adaptive)88.3%estimated ± 2.1 pp, low confidence
46Qwen3.6-27B88.3%estimated ± 1.4 pp, low confidence
47GLM-5.3-Flash88.1%estimated ± 0.5 pp, medium confidence
48Qwen3.6 Plus88.0%measured
49dots3-note Preview87.7%measured
50GPT-5.187.7%estimated ± 4.9 pp, high confidence
51Claude Opus 4.6 (Adaptive)87.5%estimated ± 4.9 pp, high confidence
52Kimi K2.687.4%measured
53Step 3.7 Flash87.2%estimated ± 2.7 pp, high confidence
54Kimi K2.5 (Reasoning)87.1%estimated ± 3.6 pp, high confidence
55GPT-5.6 Luna86.9%estimated ± 3.6 pp, high confidence
56MiMo-V2.586.8%estimated ± 2.6 pp, high confidence
57Gemini 3 Pro86.6%measured
58Grok 4.2086.6%estimated ± 1.4 pp, low confidence
59Holo2-30B-A3B86.5%estimated ± 2.4 pp, medium confidence
60Grok 4.386.3%estimated ± 3.6 pp, high confidence
61MiniMax M386.3%estimated ± 3.6 pp, high confidence
62Claude Opus 4.686.2%estimated ± 1.4 pp, low confidence
63Qwen3.5-122B-A10B86.2%measured
64Pareto 26.986.2%estimated ± 3.6 pp, high confidence
65Gemini 2.5 Pro86.0%estimated ± 4.9 pp, high confidence
66Qwen3.5-27B86.0%measured
67Claude Opus 4.886.0%estimated ± 2.1 pp, low confidence
68Qwen3.6-35B-A3B85.0%estimated ± 2.6 pp, high confidence
69Claude Sonnet 4.684.6%estimated ± 2.6 pp, high confidence
70GPT-5 (medium)84.2%estimated ± 4.9 pp, high confidence
71Gemma 4 31B84.2%estimated ± 3.6 pp, high confidence
72GPT-5 (high)84.0%estimated ± 4.9 pp, high confidence
73Qwen3.5-35B-A3B83.9%measured
74GPT-5.4 mini83.7%estimated ± 3.6 pp, high confidence
75Claude Opus 4.5 Thinking83.5%estimated ± 4.9 pp, high confidence
76GPT-5.283.0%measured
77Holo2-8B82.9%estimated ± 2.4 pp, medium confidence
78Step 5 Preview82.6%estimated ± 3.6 pp, high confidence
79Nemotron 3 Nano Omni 30B A3B82.3%estimated ± 2.4 pp, medium confidence
80MiMo-V2.6-Flash82.1%estimated ± 4.9 pp, high confidence
81Inkling82.0%estimated ± 2.1 pp, low confidence
82Holo2-4B82.0%estimated ± 2.4 pp, medium confidence
83Gemini 3.1 Flash-Lite81.9%estimated ± 2.6 pp, high confidence
84GLM-5V-Turbo81.9%estimated ± 4.9 pp, high confidence
85GPT-5.1-Codex81.7%estimated ± 4.9 pp, high confidence
86GPT-5.1-Codex-Max81.7%estimated ± 4.9 pp, high confidence
87o381.5%estimated ± 4.9 pp, high confidence
88MiMo-V2-Omni81.5%estimated ± 4.9 pp, high confidence
89Grok 481.5%estimated ± 4.9 pp, high confidence
90Claude 4.1 Opus Thinking81.5%estimated ± 4.9 pp, high confidence
91Claude 3 Haiku81.5%estimated ± 4.9 pp, medium confidence
92Claude 4 Sonnet81.5%estimated ± 4.9 pp, high confidence
93Gemini 1.5 Pro81.5%estimated ± 4.9 pp, high confidence
94Gemini 2.5 Flash81.5%estimated ± 4.9 pp, high confidence
95Gemma 3 27B81.5%estimated ± 4.9 pp, medium confidence
96Gemma 4 E2B81.5%estimated ± 4.9 pp, medium confidence
97Gemma 4 E4B81.5%estimated ± 4.9 pp, medium confidence
98GPT-4.181.5%estimated ± 4.9 pp, high confidence
99GPT-4.1 mini81.5%estimated ± 4.9 pp, high confidence
100GPT-4.1 nano81.5%estimated ± 4.9 pp, medium confidence
101GPT-4o mini81.5%estimated ± 4.9 pp, medium confidence
102Grok 4.1 Fast81.5%estimated ± 4.9 pp, medium confidence
103Grok 4.1 Fast (Reasoning)81.5%estimated ± 4.9 pp, high confidence
104Grok 4 Fast (Reasoning)81.5%estimated ± 4.9 pp, high confidence
105LFM2.5-VL-1.6B-Extract81.5%estimated ± 4.9 pp, medium confidence
106Llama 4 Maverick81.5%estimated ± 4.9 pp, high confidence
107Llama 4 Scout81.5%estimated ± 4.9 pp, high confidence
108Mistral Large 381.5%estimated ± 4.9 pp, high confidence
109Mistral Medium 381.5%estimated ± 4.9 pp, high confidence
110Mistral Medium 3.5 128B81.5%estimated ± 4.9 pp, high confidence
111Mistral Small 481.5%estimated ± 4.9 pp, high confidence
112Mistral Small 4 (Reasoning)81.5%estimated ± 4.9 pp, high confidence
113Nova Pro81.5%estimated ± 4.9 pp, medium confidence
114Phi-4 Multimodal Instruct81.5%estimated ± 4.9 pp, medium confidence
115Qwen3.5 397B (Reasoning)81.5%estimated ± 4.9 pp, high confidence
116Qwen3-Omni-30B-A3B-Instruct81.5%estimated ± 4.9 pp, high confidence
117Qwen3-Omni-30B-A3B-Thinking81.5%estimated ± 4.9 pp, high confidence
118Inkling-Small80.8%estimated ± 2.1 pp, low confidence
119Claude Sonnet 580.1%estimated ± 2.1 pp, low confidence
120Gemma 4 12B79.7%measured
121Gemma 4 26B A4B78.8%estimated ± 3.6 pp, high confidence
122Interfaze Beta74.3%estimated ± 3.6 pp, high confidence
123Claude Opus 4.574.3%measured
124Command A+66.6%estimated ± 2.6 pp, medium confidence
125GPT-5.4 nano66.5%estimated ± 3.6 pp, medium confidence
126LFM2.5-VL-3B24.3%estimated ± 3.6 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General