benchgap
Vision & documents

Video-MME (w/o subtitle) leaderboard

As of 2026-10-10, the highest measured score on Video-MME (w/o subtitle) is 87.7% by Gemini 3 Pro. 111 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Step 3.7 Flash100.0%estimated ± 3.2 pp, low confidence
2Claude Opus 5.590.8%estimated ± 6.3 pp, low confidence
3GPT-6 Astra90.2%estimated ± 6.3 pp, low confidence
4Qwen3.8 Max89.7%estimated ± 1.7 pp, low confidence
5GPT-6.1 Sol89.7%estimated ± 6.3 pp, low confidence
6Gemini 3.8 Flash89.4%estimated ± 6.3 pp, low confidence
7Gemini 3.7 Flash89.3%estimated ± 6.3 pp, low confidence
8Claude Opus 588.8%estimated ± 6.3 pp, low confidence
9Gemini 3.6 Flash87.9%estimated ± 6.3 pp, low confidence
10GPT-6 Sol87.8%estimated ± 6.3 pp, low confidence
11Gemini 3 Pro87.7%measured
12Qwen3.8 Max Preview87.7%estimated ± 6.3 pp, low confidence
13Gemini 3.1 Pro87.3%estimated ± 1.8 pp, low confidence
14Gemini 3.5 Flash87.1%estimated ± 1.8 pp, low confidence
15GPT-5.6 Sol86.8%estimated ± 1.8 pp, low confidence
16Qwen3.8-Omni-Flash86.5%estimated ± 3.0 pp, low confidence
17Grok 4.586.3%estimated ± 6.3 pp, low confidence
18dots3-note Preview86.3%estimated ± 1.7 pp, medium confidence
19Qwen3.8-Flash-Next86.2%estimated ± 3.0 pp, low confidence
20Kimi K386.2%estimated ± 1.8 pp, low confidence
21Seed 2.1 Pro86.2%estimated ± 1.8 pp, low confidence
22Qwen3.8-27B86.1%estimated ± 3.0 pp, low confidence
23GPT-5.486.0%estimated ± 1.8 pp, low confidence
24GPT-5.586.0%estimated ± 1.8 pp, low confidence
25GPT-6 Luna85.9%estimated ± 6.3 pp, low confidence
26GPT-5.285.8%measured
27GPT-5.6 Terra85.7%estimated ± 1.8 pp, medium confidence
28Apodex 1.185.6%estimated ± 6.3 pp, low confidence
29Apodex 1.1 Mini85.6%estimated ± 6.3 pp, low confidence
30Muse Spark85.6%estimated ± 1.8 pp, medium confidence
31Gemini 3.5 Flash-Lite85.5%estimated ± 6.3 pp, low confidence
32Ling 3.0 Flash VL85.5%estimated ± 6.3 pp, low confidence
33Qwen3.6 Plus85.5%estimated ± 1.0 pp, medium confidence
34Seed 2.1 Turbo85.4%estimated ± 1.8 pp, medium confidence
35Claude Opus 4.7 (Adaptive)85.4%estimated ± 6.3 pp, low confidence
36Gemini 3 Flash85.3%estimated ± 6.3 pp, low confidence
37GPT-5.3 Codex85.3%estimated ± 6.3 pp, low confidence
38Kimi K2.685.1%estimated ± 1.8 pp, medium confidence
39Kimi K2.5 (Reasoning)84.7%estimated ± 1.8 pp, medium confidence
40Claude Sonnet 584.7%estimated ± 6.3 pp, low confidence
41GPT-5.6 Luna84.6%estimated ± 1.8 pp, medium confidence
42DeepSeek V4.1 Flash84.5%estimated ± 6.3 pp, low confidence
43Grok 4.384.5%estimated ± 1.8 pp, medium confidence
44Pareto 26.984.4%estimated ± 1.8 pp, medium confidence
45MiMo-V2.584.4%estimated ± 1.8 pp, medium confidence
46Claude Opus 4.784.2%estimated ± 6.3 pp, low confidence
47Mistral Large 484.2%estimated ± 6.3 pp, low confidence
48GPT-5.2-Codex84.2%estimated ± 6.3 pp, low confidence
49Claude Opus 4.684.1%estimated ± 1.8 pp, medium confidence
50Gemma 4 31B83.9%estimated ± 1.8 pp, medium confidence
51Qwen3.7 Plus83.9%estimated ± 1.7 pp, medium confidence
52GPT-5.183.8%estimated ± 6.3 pp, low confidence
53Claude Opus 4.6 (Adaptive)83.7%estimated ± 6.3 pp, low confidence
54GPT-5.4 mini83.7%estimated ± 1.8 pp, medium confidence
55Qwen3.5-122B-A10B83.7%estimated ± 1.0 pp, medium confidence
56Qwen3.5 397B83.7%measured
57Gemini 2.5 Pro83.5%estimated ± 6.3 pp, low confidence
58Step 5 Preview83.4%estimated ± 1.8 pp, medium confidence
59GPT-5 (medium)83.2%estimated ± 6.3 pp, low confidence
60Kimi K2.583.2%measured
61GPT-5 (high)83.2%estimated ± 6.3 pp, low confidence
62Claude Opus 4.5 Thinking83.1%estimated ± 6.3 pp, low confidence
63Grok 4.2083.0%estimated ± 1.8 pp, medium confidence
64Qwen3.6-27B82.8%estimated ± 1.0 pp, medium confidence
65MiMo-V2.6-Flash82.7%estimated ± 6.3 pp, low confidence
66GLM-5V-Turbo82.6%estimated ± 6.3 pp, low confidence
67MiniMax M382.5%estimated ± 1.7 pp, medium confidence
68Qwen3.6-35B-A3B82.5%measured
69GPT-5.1-Codex82.4%estimated ± 6.3 pp, low confidence
70GPT-5.1-Codex-Max82.4%estimated ± 6.3 pp, low confidence
71Inkling-Small82.4%estimated ± 1.8 pp, medium confidence
72Muse Glimmer 30B82.4%estimated ± 1.8 pp, medium confidence
73Qwen3.5-27B82.3%estimated ± 1.0 pp, medium confidence
74Gemma 4 26B A4B82.3%estimated ± 1.8 pp, medium confidence
75Inkling82.1%estimated ± 1.8 pp, medium confidence
76Claude Sonnet 4.681.7%estimated ± 6.3 pp, low confidence
77Qwen3.5-35B-A3B81.5%estimated ± 1.0 pp, medium confidence
78o381.5%estimated ± 6.3 pp, low confidence
79MiMo-V2-Omni81.4%estimated ± 6.3 pp, low confidence
80Claude Opus 4.581.4%measured
81Grok 481.0%estimated ± 6.3 pp, low confidence
82Interfaze Beta80.9%estimated ± 1.8 pp, medium confidence
83Claude 4.1 Opus Thinking80.7%estimated ± 6.3 pp, low confidence
84Gemini 2.5 Flash80.0%estimated ± 6.3 pp, low confidence
85Gemma 4 12B79.9%estimated ± 1.8 pp, low confidence
86Mistral Medium 3.5 128B79.8%estimated ± 6.3 pp, low confidence
87Grok 4.1 Fast (Reasoning)79.4%estimated ± 6.3 pp, low confidence
88Claude 4 Sonnet79.2%estimated ± 6.3 pp, low confidence
89Llama 4 Maverick79.1%estimated ± 6.3 pp, low confidence
90Grok 4 Fast (Reasoning)79.1%estimated ± 6.3 pp, low confidence
91GPT-4.178.9%estimated ± 6.3 pp, low confidence
92Qwen3-Omni-30B-A3B-Thinking78.7%estimated ± 6.3 pp, low confidence
93GPT-4.1 mini78.5%estimated ± 6.3 pp, low confidence
94GPT-5.4 nano78.3%estimated ± 1.8 pp, low confidence
95Mistral Small 478.2%estimated ± 6.3 pp, low confidence
96Mistral Small 4 (Reasoning)78.2%estimated ± 6.3 pp, low confidence
97Mistral Large 378.0%estimated ± 6.3 pp, low confidence
98Qwen3-Omni-30B-A3B-Instruct78.0%estimated ± 6.3 pp, low confidence
99Gemini 1.5 Pro77.9%estimated ± 6.3 pp, low confidence
100Mistral Medium 377.7%estimated ± 6.3 pp, low confidence
101Llama 4 Scout77.6%estimated ± 6.3 pp, low confidence
102Qwen3.5 397B (Reasoning)77.6%estimated ± 6.3 pp, low confidence
103Grok 4.1 Fast77.2%estimated ± 6.3 pp, low confidence
104Gemma 3 27B77.2%estimated ± 6.3 pp, low confidence
105Nova Pro77.0%estimated ± 6.3 pp, low confidence
106GPT-4o mini76.9%estimated ± 6.3 pp, low confidence
107GPT-4.1 nano76.9%estimated ± 6.3 pp, low confidence
108Claude 3 Haiku76.7%estimated ± 6.3 pp, low confidence
109LFM2.5-VL-1.6B-Extract76.7%estimated ± 6.3 pp, low confidence
110Phi-4 Multimodal Instruct76.7%estimated ± 6.3 pp, low confidence
111Command A+76.1%estimated ± 1.0 pp, medium confidence
112Nemotron 3 Nano Omni 30B A3B72.2%measured
113Gemma 4 E4B70.7%estimated ± 1.8 pp, low confidence
114Qwen2.5-VL-32B68.8%estimated ± 1.8 pp, low confidence
115Gemma 4 E2B65.5%estimated ± 1.8 pp, low confidence
116LFM2.5-VL-3B52.9%estimated ± 1.0 pp, low confidence
117ZAYA1-VL-8B50.8%estimated ± 1.0 pp, low confidence
118LFM2.5-VL-450M39.2%estimated ± 1.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General