benchgap
Vision & documents

MLVU (M-Avg) leaderboard

As of 2026-10-10, the highest measured score on MLVU (M-Avg) is 90.8% by Qwen3.8 Max. 64 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.8 Max90.8%measured
2GPT-6 Astra90.4%estimated ± 2.9 pp, low confidence
3Claude Mythos 590.1%estimated ± 1.5 pp, high confidence
4Kimi K389.4%estimated ± 1.5 pp, high confidence
5Claude Opus 4.7 (Adaptive)89.3%estimated ± 1.5 pp, high confidence
6Claude Opus 4.888.9%estimated ± 1.5 pp, high confidence
7GLM-5.3-Flash88.8%estimated ± 1.5 pp, high confidence
8Qwen3.8-Flash-Next88.7%estimated ± 1.5 pp, medium confidence
9Gemini 3.7 Flash88.5%estimated ± 1.5 pp, high confidence
10Muse Spark 1.188.4%estimated ± 1.5 pp, high confidence
11Claude Sonnet 588.4%estimated ± 1.5 pp, high confidence
12Qwen3.8-Omni-Flash88.2%estimated ± 1.5 pp, high confidence
13GPT-5.6 Sol88.0%estimated ± 2.7 pp, low confidence
14Sakana Fugu-Ultra87.8%estimated ± 1.5 pp, high confidence
15Muse Spark87.8%estimated ± 1.5 pp, high confidence
16Seed 2.1 Pro87.4%estimated ± 1.5 pp, high confidence
17Qwen3.7 Plus87.4%measured
18Sakana Fugu87.3%estimated ± 1.5 pp, high confidence
19GPT-5.587.3%estimated ± 2.7 pp, medium confidence
20Qwen3.8-27B87.2%estimated ± 1.5 pp, high confidence
21GPT-5.6 Terra87.1%estimated ± 2.7 pp, medium confidence
22Gemini 3.5 Flash87.0%estimated ± 1.5 pp, high confidence
23Qwen3.5 397B86.7%measured
24Qwen3.6-27B86.6%measured
25GPT-5.486.6%estimated ± 1.5 pp, high confidence
26Seed 2.1 Turbo86.5%estimated ± 1.5 pp, high confidence
27Inkling86.3%estimated ± 1.5 pp, high confidence
28dots3-note Preview86.2%estimated ± 1.6 pp, medium confidence
29Qwen3.6-35B-A3B86.2%measured
30Qwen3.6 Plus86.2%estimated ± 1.5 pp, high confidence
31GPT-5.6 Luna86.1%estimated ± 2.7 pp, medium confidence
32Holo2-235B-A22B86.1%estimated ± 2.9 pp, medium confidence
33Inkling-Small86.1%estimated ± 1.5 pp, high confidence
34Step 3.7 Flash86.1%estimated ± 2.2 pp, medium confidence
35Grok 4.386.0%estimated ± 2.7 pp, medium confidence
36MiMo-V2.586.0%estimated ± 1.5 pp, high confidence
37Pareto 26.986.0%estimated ± 2.7 pp, medium confidence
38Nemotron 3 Nano Omni 30B A3B85.8%estimated ± 0.4 pp, medium confidence
39Kimi K2.685.8%estimated ± 1.5 pp, high confidence
40Gemini 3.1 Pro85.7%estimated ± 1.5 pp, high confidence
41GPT-5.285.6%measured
42Gemma 4 31B85.5%estimated ± 2.7 pp, medium confidence
43Qwen3.5-27B85.5%estimated ± 1.5 pp, high confidence
44GPT-5.4 mini85.4%estimated ± 2.7 pp, medium confidence
45MiniMax M385.3%estimated ± 2.1 pp, high confidence
46Muse Glimmer 30B85.3%estimated ± 1.5 pp, high confidence
47Holo2-30B-A3B85.3%estimated ± 2.9 pp, medium confidence
48Step 5 Preview85.1%estimated ± 2.7 pp, medium confidence
49Kimi K2.585.0%measured
50Kimi K2.5 (Reasoning)84.8%estimated ± 1.5 pp, high confidence
51Qwen3.5-35B-A3B84.8%estimated ± 1.5 pp, high confidence
52Claude Sonnet 4.684.8%estimated ± 1.5 pp, high confidence
53LFM2.5-VL-3B84.8%estimated ± 0.4 pp, medium confidence
54Qwen3.5-122B-A10B84.7%estimated ± 1.5 pp, high confidence
55Gemma 4 26B A4B84.2%estimated ± 2.7 pp, medium confidence
56Holo2-8B83.9%estimated ± 2.9 pp, medium confidence
57Ternary Bonsai 2 27B83.8%estimated ± 1.5 pp, high confidence
58Holo2-4B83.5%estimated ± 2.9 pp, medium confidence
59Gemini 3.1 Flash-Lite83.4%estimated ± 1.5 pp, high confidence
60ZAYA1-VL-8B83.2%estimated ± 0.4 pp, medium confidence
61Gemma 4 12B83.1%estimated ± 1.6 pp, medium confidence
62Gemini 3 Pro83.0%measured
63Claude Opus 4.683.0%estimated ± 2.5 pp, medium confidence
64Interfaze Beta82.3%estimated ± 0.4 pp, low confidence
65Claude Opus 4.581.7%measured
66GPT-5.4 nano80.5%estimated ± 2.7 pp, low confidence
67Grok 4.2079.4%estimated ± 1.5 pp, medium confidence
68Command A+76.7%estimated ± 1.5 pp, medium confidence
69Gemma 4 E4B72.6%estimated ± 2.7 pp, low confidence
70North Micro Vision Instruct71.8%estimated ± 1.5 pp, medium confidence
71Qwen2.5-VL-32B70.5%estimated ± 2.7 pp, low confidence
72LFM2.5-VL-450M68.9%estimated ± 1.5 pp, medium confidence
73Gemma 4 E2B66.6%estimated ± 2.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General