benchgap
Vision & documents

MMMU-Pro leaderboard

As of 2026-10-07, the highest measured score on MMMU-Pro is 88.0% by Claude Opus 5.5 (max with fallback). 13 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Claude Opus 5.5 (max with fallback)88.0%measured
2GPT-6 Astra (max)87.0%measured
3Claude Opus 5.5 (xhigh with fallback)87.0%measured
4GPT-6.1 Sol (high)86.7%estimated ± 2.6 pp, high confidence
5GPT-6.1 Sol (xhigh)86.7%estimated ± 2.6 pp, high confidence
6GPT-6 Astra (medium)86.4%estimated ± 2.6 pp, high confidence
7GPT-6 Astra (high)86.0%measured
8Gemini 3.8 Flash (high)86.0%measured
9GPT-6 Astra (xhigh)86.0%measured
10Claude Opus 5.5 (high with fallback)86.0%measured
11GPT-6.1 Sol (max)86.0%measured
12Claude Opus 5.5 (medium with fallback)86.0%measured
13Muse Spark 1.3 (max)85.0%estimated ± 2.6 pp, high confidence
14Claude Fable 5.1 (max with fallback)84.8%estimated ± 2.6 pp, high confidence
15Claude Sonnet 5.5 (max with fallback)84.6%estimated ± 2.6 pp, high confidence
16Qwen3.8 Max (0902)83.0%measured
17GPT-6 Sol (max)83.0%measured
18Gemini 4 Argon (high)81.6%estimated ± 2.6 pp, high confidence
19Kimi K3 (max)81.0%measured
20Grok 4.7 (xhigh)80.0%estimated ± 2.6 pp, high confidence
21GPT-6 Luna (max)80.0%measured
22MiMo-V2.6-Pro79.3%estimated ± 2.6 pp, high confidence
23MiniMax-M379.0%measured
24DeepSeek V4.1 Flash (max)77.0%measured
25GLM-5.3-Flash76.6%estimated ± 2.6 pp, high confidence
26Qwen3.8 27B (xhigh)76.0%measured
27Step 5 Preview76.0%measured
28Mistral Large 4 Preview76.0%measured
29GLM-5.3 (max)75.3%estimated ± 2.6 pp, high confidence
30K2 Horizon 375B A23B75.1%estimated ± 2.6 pp, medium confidence
31Nemotron 3 Ultra75.1%estimated ± 2.6 pp, medium confidence
32Muse Glimmer (high)74.0%measured
33Inkling (xhigh)73.0%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents