benchgap
Vision & documents

VideoMMMU leaderboard

As of 2026-10-07, the highest measured score on VideoMMMU is 88.7% by Qwen3.8 Max. 55 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra90.6%estimated ± 1.7 pp, low confidence
2Gemini 3.1 Pro89.3%estimated ± 1.0 pp, medium confidence
3Gemini 3.5 Flash89.2%estimated ± 1.0 pp, medium confidence
4GPT-5.6 Sol89.0%estimated ± 1.0 pp, medium confidence
5Qwen3.8 Max88.7%measured
6Kimi K388.2%estimated ± 1.0 pp, high confidence
7Seed 2.1 Pro88.2%estimated ± 1.0 pp, high confidence
8Claude Mythos 588.2%estimated ± 1.5 pp, high confidence
9GPT-5.487.8%estimated ± 1.0 pp, high confidence
10GPT-5.587.8%estimated ± 1.0 pp, high confidence
11Gemini 3 Pro87.6%measured
12Qwen3.8-Omni-Flash87.6%estimated ± 1.5 pp, high confidence
13Claude Opus 4.7 (Adaptive)87.4%estimated ± 1.5 pp, high confidence
14Qwen3.8-Flash-Next87.3%estimated ± 1.5 pp, high confidence
15GPT-5.6 Terra87.3%estimated ± 1.0 pp, high confidence
16Qwen3.8-27B87.2%estimated ± 1.5 pp, high confidence
17Claude Opus 4.887.1%estimated ± 1.5 pp, high confidence
18Step 3.7 Flash87.0%estimated ± 2.2 pp, low confidence
19GLM-5.3-Flash87.0%estimated ± 1.5 pp, high confidence
20Muse Spark87.0%estimated ± 1.0 pp, high confidence
21Gemini 3.7 Flash86.8%estimated ± 1.5 pp, high confidence
22dots3-note Preview86.8%measured
23Muse Spark 1.186.8%estimated ± 1.5 pp, high confidence
24Claude Sonnet 586.7%estimated ± 1.5 pp, high confidence
25Seed 2.1 Turbo86.6%estimated ± 1.0 pp, high confidence
26Kimi K2.586.6%measured
27Sakana Fugu-Ultra86.3%estimated ± 1.5 pp, high confidence
28GPT-5.286.0%estimated ± 1.0 pp, high confidence
29Sakana Fugu86.0%estimated ± 1.5 pp, high confidence
30Kimi K2.685.9%estimated ± 1.0 pp, high confidence
31Qwen3.5-27B85.6%estimated ± 1.9 pp, low confidence
32Holo2-235B-A22B85.5%estimated ± 1.7 pp, medium confidence
33Qwen3.7 Plus85.4%measured
34Qwen3.5-35B-A3B85.2%estimated ± 1.9 pp, low confidence
35Kimi K2.5 (Reasoning)85.1%estimated ± 1.0 pp, high confidence
36GPT-5.6 Luna85.1%estimated ± 1.0 pp, high confidence
37Holo2-30B-A3B85.0%estimated ± 1.7 pp, medium confidence
38Grok 4.384.9%estimated ± 1.0 pp, high confidence
39Pareto 26.984.8%estimated ± 1.0 pp, high confidence
40MiMo-V2.584.8%estimated ± 1.0 pp, high confidence
41Qwen3.5 397B84.7%measured
42MiniMax M384.6%measured
43Claude Sonnet 4.684.6%estimated ± 1.5 pp, high confidence
44Qwen3.5-122B-A10B84.5%estimated ± 1.5 pp, high confidence
45Claude Opus 4.684.5%estimated ± 1.0 pp, high confidence
46Holo2-8B84.4%estimated ± 1.7 pp, medium confidence
47Gemma 4 31B84.4%estimated ± 1.0 pp, high confidence
48Nemotron 3 Nano Omni 30B A3B84.4%estimated ± 1.5 pp, high confidence
49Claude Opus 4.584.4%measured
50Qwen3.6-27B84.4%measured
51GPT-5.4 mini84.4%estimated ± 1.0 pp, high confidence
52Holo2-4B84.4%estimated ± 1.7 pp, medium confidence
53Step 5 Preview84.3%estimated ± 1.0 pp, high confidence
54Grok 4.2084.2%estimated ± 1.0 pp, high confidence
55Inkling-Small84.1%estimated ± 1.0 pp, high confidence
56Muse Glimmer 30B84.1%estimated ± 1.0 pp, high confidence
57Gemma 4 26B A4B84.1%estimated ± 1.0 pp, high confidence
58Inkling84.1%estimated ± 1.0 pp, high confidence
59Interfaze Beta84.1%estimated ± 1.0 pp, high confidence
60Gemma 4 12B84.1%estimated ± 1.0 pp, medium confidence
61Command A+84.1%estimated ± 1.0 pp, medium confidence
62GPT-5.4 nano84.1%estimated ± 1.0 pp, medium confidence
63LFM2.5-VL-3B84.1%estimated ± 1.0 pp, medium confidence
64Gemini 3.1 Flash-Lite84.0%estimated ± 1.5 pp, high confidence
65Qwen3.6 Plus84.0%measured
66Qwen3.6-35B-A3B83.7%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General