benchgap
Vision & documents

Video-MME (with subtitle) leaderboard

As of 2026-10-07, the highest measured score on Video-MME (with subtitle) is 90.4% by Qwen3.8 Max. 52 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.8 Max90.4%measured
2Claude Mythos 590.1%estimated ± 0.9 pp, medium confidence
3GPT-5.6 Sol90.0%estimated ± 1.8 pp, low confidence
4Qwen3.8-Omni-Flash89.7%estimated ± 0.9 pp, medium confidence
5Kimi K389.7%estimated ± 0.9 pp, medium confidence
6Claude Opus 4.7 (Adaptive)89.6%estimated ± 0.9 pp, medium confidence
7Qwen3.8-Flash-Next89.5%estimated ± 0.9 pp, medium confidence
8Qwen3.8-27B89.5%estimated ± 0.9 pp, medium confidence
9Claude Opus 4.889.4%estimated ± 0.9 pp, medium confidence
10GLM-5.3-Flash89.3%estimated ± 0.9 pp, medium confidence
11Gemini 3.7 Flash89.2%estimated ± 0.9 pp, medium confidence
12Muse Spark 1.189.1%estimated ± 0.9 pp, medium confidence
13GPT-5.589.1%estimated ± 1.8 pp, medium confidence
14Claude Sonnet 589.1%estimated ± 0.9 pp, medium confidence
15GPT-5.6 Terra88.8%estimated ± 1.8 pp, medium confidence
16dots3-note Preview88.8%estimated ± 1.2 pp, medium confidence
17Sakana Fugu-Ultra88.7%estimated ± 0.9 pp, medium confidence
18Muse Spark88.7%estimated ± 0.9 pp, medium confidence
19Kimi K2.588.6%estimated ± 1.2 pp, medium confidence
20Seed 2.1 Pro88.5%estimated ± 0.9 pp, medium confidence
21Sakana Fugu88.4%estimated ± 0.9 pp, medium confidence
22Gemini 3.5 Flash88.3%estimated ± 0.9 pp, medium confidence
23Qwen3.7 Plus88.0%measured
24GPT-5.488.0%estimated ± 0.9 pp, medium confidence
25Seed 2.1 Turbo87.9%estimated ± 0.9 pp, medium confidence
26GPT-5.287.8%estimated ± 0.9 pp, medium confidence
27Inkling87.8%estimated ± 0.9 pp, medium confidence
28Kimi K2.5 (Reasoning)87.8%estimated ± 1.8 pp, medium confidence
29GPT-5.6 Luna87.7%estimated ± 1.8 pp, medium confidence
30Qwen3.6 Plus87.7%estimated ± 0.9 pp, medium confidence
31MiMo-V2.587.7%measured
32Qwen3.6-27B87.7%measured
33Gemini 3 Pro87.7%estimated ± 0.9 pp, medium confidence
34Inkling-Small87.7%estimated ± 0.9 pp, medium confidence
35Grok 4.387.6%estimated ± 1.8 pp, medium confidence
36Qwen3.5 397B87.6%estimated ± 0.9 pp, medium confidence
37Pareto 26.987.6%estimated ± 1.8 pp, medium confidence
38Kimi K2.687.5%estimated ± 0.9 pp, medium confidence
39Gemini 3.1 Pro87.4%estimated ± 0.9 pp, medium confidence
40Claude Opus 4.687.2%estimated ± 1.8 pp, medium confidence
41Muse Glimmer 30B87.2%estimated ± 0.9 pp, medium confidence
42Gemma 4 31B87.1%estimated ± 1.8 pp, medium confidence
43GPT-5.4 mini86.9%estimated ± 1.8 pp, medium confidence
44Claude Sonnet 4.686.9%estimated ± 0.9 pp, low confidence
45Qwen3.5-122B-A10B86.8%estimated ± 0.9 pp, low confidence
46Step 5 Preview86.7%estimated ± 1.8 pp, medium confidence
47Nemotron 3 Nano Omni 30B A3B86.6%estimated ± 0.9 pp, low confidence
48Qwen3.6-35B-A3B86.6%measured
49Gemini 3.1 Flash-Lite86.0%estimated ± 0.9 pp, low confidence
50Gemma 4 26B A4B85.8%estimated ± 1.8 pp, low confidence
51MiniMax M385.4%measured
52Claude Opus 4.585.1%estimated ± 0.9 pp, low confidence
53Interfaze Beta84.8%estimated ± 1.8 pp, low confidence
54Gemma 4 12B84.2%estimated ± 1.8 pp, low confidence
55Grok 4.2083.5%estimated ± 0.9 pp, low confidence
56GPT-5.4 nano83.4%estimated ± 1.8 pp, low confidence
57Command A+81.9%estimated ± 0.9 pp, low confidence
58LFM2.5-VL-3B80.4%estimated ± 1.8 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General