benchgap
Vision & documents

OmniDocBench 1.5 leaderboard

As of 2026-10-07, the highest measured score on OmniDocBench 1.5 is 92.1% by Qwen3.8 Max. 116 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Mythos 592.7%estimated ± 7.4 pp, low confidence
2Qwen3.8 Max92.1%measured
3Claude Opus 592.0%estimated ± 6.4 pp, low confidence
4Claude Opus 5.592.0%estimated ± 6.4 pp, low confidence
5Gemini 3.6 Flash92.0%estimated ± 6.4 pp, low confidence
6Gemini 3.7 Flash92.0%estimated ± 6.4 pp, low confidence
7Gemini 3.8 Flash92.0%estimated ± 6.4 pp, low confidence
8GPT-6.1 Sol92.0%estimated ± 6.4 pp, low confidence
9GPT-6 Astra92.0%estimated ± 6.4 pp, low confidence
10GPT-6 Sol92.0%estimated ± 6.4 pp, low confidence
11Qwen3.8 Max Preview92.0%estimated ± 6.4 pp, low confidence
12Grok 4.592.0%estimated ± 6.4 pp, low confidence
13Qwen3.8-Flash-Next92.0%estimated ± 6.4 pp, low confidence
14GPT-6 Luna92.0%estimated ± 6.4 pp, low confidence
15Apodex 1.192.0%estimated ± 6.4 pp, low confidence
16Apodex 1.1 Mini92.0%estimated ± 6.4 pp, low confidence
17Gemini 3.5 Flash-Lite92.0%estimated ± 6.4 pp, low confidence
18Ling 3.0 Flash VL92.0%estimated ± 6.4 pp, low confidence
19Claude Opus 4.7 (Adaptive)92.0%estimated ± 6.4 pp, low confidence
20Gemini 3 Flash92.0%estimated ± 6.4 pp, low confidence
21GPT-5.3 Codex92.0%estimated ± 6.4 pp, low confidence
22Claude Sonnet 592.0%estimated ± 6.4 pp, low confidence
23DeepSeek V4.1 Flash91.9%estimated ± 6.4 pp, low confidence
24Claude Opus 4.791.7%estimated ± 6.4 pp, low confidence
25Mistral Large 491.7%estimated ± 6.4 pp, low confidence
26Gemini 3.1 Pro91.7%estimated ± 5.9 pp, low confidence
27Gemini 3.5 Flash91.7%estimated ± 5.9 pp, low confidence
28GPT-5.6 Sol91.7%estimated ± 5.9 pp, low confidence
29Kimi K391.7%estimated ± 5.9 pp, low confidence
30Seed 2.1 Pro91.7%estimated ± 5.9 pp, low confidence
31Gemini 3 Pro91.7%estimated ± 5.9 pp, low confidence
32GPT-5.491.7%estimated ± 5.9 pp, low confidence
33GPT-5.591.7%estimated ± 5.9 pp, low confidence
34GPT-5.6 Terra91.7%estimated ± 5.9 pp, low confidence
35Muse Spark91.7%estimated ± 5.9 pp, low confidence
36Seed 2.1 Turbo91.7%estimated ± 5.9 pp, low confidence
37GPT-5.291.7%estimated ± 5.9 pp, low confidence
38Kimi K2.691.7%estimated ± 5.9 pp, low confidence
39dots3-note Preview91.7%estimated ± 5.9 pp, low confidence
40Qwen3.5 397B91.7%estimated ± 5.9 pp, low confidence
41Qwen3.6 Plus91.7%estimated ± 5.9 pp, low confidence
42Kimi K2.591.7%estimated ± 5.9 pp, low confidence
43Kimi K2.5 (Reasoning)91.7%estimated ± 5.9 pp, low confidence
44GPT-5.6 Luna91.7%estimated ± 5.9 pp, low confidence
45Grok 4.391.7%estimated ± 5.9 pp, low confidence
46Pareto 26.991.7%estimated ± 5.9 pp, low confidence
47MiMo-V2.591.7%estimated ± 5.9 pp, low confidence
48GPT-5.2-Codex91.7%estimated ± 6.4 pp, low confidence
49Claude Opus 4.691.7%estimated ± 5.9 pp, low confidence
50Qwen3.8-Omni-Flash91.6%estimated ± 7.4 pp, low confidence
51Gemma 4 31B91.6%estimated ± 5.9 pp, low confidence
52MiniMax M391.6%measured
53GPT-5.4 mini91.5%estimated ± 5.9 pp, low confidence
54Qwen3.7 Plus91.4%measured
55Step 5 Preview91.2%estimated ± 5.9 pp, low confidence
56Qwen3.8-27B91.1%measured
57Qwen3.6-27B91.0%estimated ± 5.9 pp, low confidence
58Claude Opus 4.890.8%estimated ± 7.4 pp, low confidence
59GLM-5.3-Flash90.5%estimated ± 7.4 pp, low confidence
60GPT-5.190.3%estimated ± 6.4 pp, low confidence
61Claude Opus 4.6 (Adaptive)90.0%estimated ± 6.4 pp, low confidence
62Muse Spark 1.189.9%estimated ± 7.4 pp, low confidence
63Qwen3.6-35B-A3B89.9%measured
64Step 3.7 Flash89.6%estimated ± 6.4 pp, low confidence
65Grok 4.2089.5%estimated ± 5.9 pp, low confidence
66Sakana Fugu-Ultra88.9%estimated ± 7.4 pp, low confidence
67Sakana Fugu88.1%estimated ± 7.4 pp, low confidence
68Qwen3.5-122B-A10B87.6%estimated ± 6.4 pp, low confidence
69Qwen3.5-27B87.6%estimated ± 6.4 pp, low confidence
70Gemini 2.5 Pro86.7%estimated ± 6.4 pp, low confidence
71Gemini 3.1 Flash-Lite80.7%estimated ± 7.4 pp, low confidence
72GPT-5 (medium)76.5%estimated ± 6.4 pp, low confidence
73Inkling-Small75.8%estimated ± 5.9 pp, low confidence
74Muse Glimmer 30B75.8%measured
75GPT-5 (high)73.8%estimated ± 6.4 pp, low confidence
76Gemma 4 26B A4B70.5%estimated ± 5.9 pp, low confidence
77Claude Opus 4.5 Thinking67.2%estimated ± 6.4 pp, low confidence
78Inkling60.5%estimated ± 5.9 pp, low confidence
79MiMo-V2.6-Flash28.5%estimated ± 6.4 pp, low confidence
80GLM-5V-Turbo18.2%estimated ± 6.4 pp, low confidence
81Qwen3.5-35B-A3B15.4%estimated ± 6.4 pp, low confidence
82GPT-5.1-Codex10.9%estimated ± 6.4 pp, low confidence
83GPT-5.1-Codex-Max10.9%estimated ± 6.4 pp, low confidence
84Interfaze Beta2.3%estimated ± 5.9 pp, low confidence
85Claude Opus 4.50.9%estimated ± 5.9 pp, low confidence
86Claude Sonnet 4.60.3%estimated ± 6.4 pp, low confidence
87o30.1%estimated ± 6.4 pp, low confidence
88MiMo-V2-Omni0.1%estimated ± 6.4 pp, low confidence
89Gemma 4 12B0.1%estimated ± 5.9 pp, low confidence
90Grok 40.0%estimated ± 6.4 pp, low confidence
91Claude 4.1 Opus Thinking0.0%estimated ± 6.4 pp, low confidence
92GPT-5.4 nano0.0%estimated ± 5.9 pp, low confidence
93Command A+0.0%estimated ± 5.9 pp, low confidence
94LFM2.5-VL-3B0.0%estimated ± 5.9 pp, low confidence
95Claude 3 Haiku0.0%estimated ± 6.4 pp, low confidence
96Claude 4 Sonnet0.0%estimated ± 6.4 pp, low confidence
97Gemini 1.5 Pro0.0%estimated ± 6.4 pp, low confidence
98Gemini 2.5 Flash0.0%estimated ± 6.4 pp, low confidence
99Gemma 3 27B0.0%estimated ± 6.4 pp, low confidence
100Gemma 4 E2B0.0%estimated ± 6.4 pp, low confidence
101Gemma 4 E4B0.0%estimated ± 6.4 pp, low confidence
102GPT-4.10.0%estimated ± 6.4 pp, low confidence
103GPT-4.1 mini0.0%estimated ± 6.4 pp, low confidence
104GPT-4.1 nano0.0%estimated ± 6.4 pp, low confidence
105GPT-4o mini0.0%estimated ± 6.4 pp, low confidence
106Grok 4.1 Fast0.0%estimated ± 6.4 pp, low confidence
107Grok 4.1 Fast (Reasoning)0.0%estimated ± 6.4 pp, low confidence
108Grok 4 Fast (Reasoning)0.0%estimated ± 6.4 pp, low confidence
109LFM2.5-VL-1.6B-Extract0.0%estimated ± 6.4 pp, low confidence
110Llama 4 Maverick0.0%estimated ± 6.4 pp, low confidence
111Llama 4 Scout0.0%estimated ± 6.4 pp, low confidence
112Mistral Large 30.0%estimated ± 6.4 pp, low confidence
113Mistral Medium 30.0%estimated ± 6.4 pp, low confidence
114Mistral Medium 3.5 128B0.0%estimated ± 6.4 pp, low confidence
115Mistral Small 40.0%estimated ± 6.4 pp, low confidence
116Mistral Small 4 (Reasoning)0.0%estimated ± 6.4 pp, low confidence
117Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 6.4 pp, low confidence
118Nova Pro0.0%estimated ± 6.4 pp, low confidence
119Phi-4 Multimodal Instruct0.0%estimated ± 6.4 pp, low confidence
120Qwen3.5 397B (Reasoning)0.0%estimated ± 6.4 pp, low confidence
121Qwen3-Omni-30B-A3B-Instruct0.0%estimated ± 6.4 pp, low confidence
122Qwen3-Omni-30B-A3B-Thinking0.0%estimated ± 6.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General