benchgap
Vision & documents

MathVision w/ Python leaderboard

As of 2026-10-07, the highest measured score on MathVision w/ Python is 97.8% by Kimi K3. 39 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Kimi K397.8%measured
2Qwen3.8 Max97.7%measured
3Claude Mythos 597.7%estimated ± 0.9 pp, medium confidence
4Gemini 3.8 Flash97.0%estimated ± 1.3 pp, low confidence
5Claude Opus 4.7 (Adaptive)96.6%estimated ± 0.9 pp, medium confidence
6Seed 2.1 Pro96.5%estimated ± 0.5 pp, medium confidence
7Qwen3.8-Omni-Flash96.2%measured
8Qwen3.8-Flash-Next95.7%measured
9Qwen3.7 Plus95.3%estimated ± 0.5 pp, medium confidence
10Seed 2.1 Turbo95.2%estimated ± 0.5 pp, medium confidence
11Qwen3.8-27B94.6%measured
12Qwen3.5 397B94.3%estimated ± 0.5 pp, low confidence
13Qwen3.6 Plus94.0%estimated ± 0.5 pp, low confidence
14dots3-note Preview93.8%estimated ± 0.5 pp, low confidence
15Kimi K2.693.6%estimated ± 0.5 pp, low confidence
16Claude Opus 4.893.3%estimated ± 0.9 pp, low confidence
17Gemini 3 Pro93.2%estimated ± 0.5 pp, low confidence
18Qwen3.5-122B-A10B93.0%estimated ± 0.5 pp, low confidence
19Qwen3.5-27B92.8%estimated ± 0.5 pp, low confidence
20Qwen3.5-35B-A3B91.6%estimated ± 0.5 pp, low confidence
21GPT-5.291.1%estimated ± 0.5 pp, low confidence
22GLM-5.3-Flash90.4%estimated ± 0.9 pp, low confidence
23Gemma 4 12B89.1%estimated ± 0.5 pp, low confidence
24Claude Opus 4.585.6%estimated ± 0.5 pp, low confidence
25Gemini 3.7 Flash85.5%estimated ± 0.9 pp, low confidence
26Muse Spark 1.183.5%estimated ± 0.9 pp, low confidence
27Claude Sonnet 582.9%estimated ± 0.9 pp, low confidence
28Sakana Fugu-Ultra77.6%estimated ± 0.9 pp, low confidence
29Muse Spark77.4%estimated ± 0.9 pp, low confidence
30Sakana Fugu76.9%estimated ± 0.9 pp, low confidence
31Gemini 3.5 Flash76.8%estimated ± 0.9 pp, low confidence
32GPT-5.476.8%estimated ± 0.9 pp, low confidence
33Inkling76.8%estimated ± 0.9 pp, low confidence
34Inkling-Small76.8%estimated ± 0.9 pp, low confidence
35MiMo-V2.576.8%estimated ± 0.9 pp, low confidence
36Gemini 3.1 Pro76.8%estimated ± 0.9 pp, low confidence
37Muse Glimmer 30B76.8%estimated ± 0.9 pp, low confidence
38Qwen3.6-27B76.8%estimated ± 0.9 pp, low confidence
39Claude Sonnet 4.676.8%estimated ± 0.9 pp, low confidence
40Command A+76.8%estimated ± 0.9 pp, low confidence
41Gemini 3.1 Flash-Lite76.8%estimated ± 0.9 pp, low confidence
42Grok 4.2076.8%estimated ± 0.9 pp, low confidence
43Nemotron 3 Nano Omni 30B A3B76.8%estimated ± 0.9 pp, low confidence
44Qwen3.6-35B-A3B76.8%estimated ± 0.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General