Vision & documents
MMMU-Pro leaderboard
As of 2026-10-07, the highest measured score on MMMU-Pro is 88.0% by Claude Opus 5.5 (max with fallback). 13 more models have estimated scores, calibrated from the benchmarks they were measured on.
Measured scores: artificial-analysis.
| # | Model | Score | Source |
|---|---|---|---|
| 1 | Claude Opus 5.5 (max with fallback) | 88.0% | measured |
| 2 | GPT-6 Astra (max) | 87.0% | measured |
| 3 | Claude Opus 5.5 (xhigh with fallback) | 87.0% | measured |
| 4 | GPT-6.1 Sol (high) | 86.7% | estimated ± 2.6 pp, high confidence |
| 5 | GPT-6.1 Sol (xhigh) | 86.7% | estimated ± 2.6 pp, high confidence |
| 6 | GPT-6 Astra (medium) | 86.4% | estimated ± 2.6 pp, high confidence |
| 7 | GPT-6 Astra (high) | 86.0% | measured |
| 8 | Gemini 3.8 Flash (high) | 86.0% | measured |
| 9 | GPT-6 Astra (xhigh) | 86.0% | measured |
| 10 | Claude Opus 5.5 (high with fallback) | 86.0% | measured |
| 11 | GPT-6.1 Sol (max) | 86.0% | measured |
| 12 | Claude Opus 5.5 (medium with fallback) | 86.0% | measured |
| 13 | Muse Spark 1.3 (max) | 85.0% | estimated ± 2.6 pp, high confidence |
| 14 | Claude Fable 5.1 (max with fallback) | 84.8% | estimated ± 2.6 pp, high confidence |
| 15 | Claude Sonnet 5.5 (max with fallback) | 84.6% | estimated ± 2.6 pp, high confidence |
| 16 | Qwen3.8 Max (0902) | 83.0% | measured |
| 17 | GPT-6 Sol (max) | 83.0% | measured |
| 18 | Gemini 4 Argon (high) | 81.6% | estimated ± 2.6 pp, high confidence |
| 19 | Kimi K3 (max) | 81.0% | measured |
| 20 | Grok 4.7 (xhigh) | 80.0% | estimated ± 2.6 pp, high confidence |
| 21 | GPT-6 Luna (max) | 80.0% | measured |
| 22 | MiMo-V2.6-Pro | 79.3% | estimated ± 2.6 pp, high confidence |
| 23 | MiniMax-M3 | 79.0% | measured |
| 24 | DeepSeek V4.1 Flash (max) | 77.0% | measured |
| 25 | GLM-5.3-Flash | 76.6% | estimated ± 2.6 pp, high confidence |
| 26 | Qwen3.8 27B (xhigh) | 76.0% | measured |
| 27 | Step 5 Preview | 76.0% | measured |
| 28 | Mistral Large 4 Preview | 76.0% | measured |
| 29 | GLM-5.3 (max) | 75.3% | estimated ± 2.6 pp, high confidence |
| 30 | K2 Horizon 375B A23B | 75.1% | estimated ± 2.6 pp, medium confidence |
| 31 | Nemotron 3 Ultra | 75.1% | estimated ± 2.6 pp, medium confidence |
| 32 | Muse Glimmer (high) | 74.0% | measured |
| 33 | Inkling (xhigh) | 73.0% | measured |
All leaderboards
Agentic · terminal
Agentic · tools
- AA-AnalystAgent
- APEX-Agents
- AutomationBench
- EnterpriseOps-Gym
- Harvey LAB
- ITBench SRE
- τ³-Bench Banking
- τ²-Bench Telecom