benchgap
Vision & documents

OfficeQA Pro leaderboard

As of 2026-10-07, the highest measured score on OfficeQA Pro is 67.7% by Claude Opus 5.5. 96 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra67.7%estimated ± 7.2 pp, medium confidence
2GPT-6.1 Sol67.7%estimated ± 7.2 pp, medium confidence
3Gemini 3.8 Flash67.7%estimated ± 7.2 pp, medium confidence
4Gemini 3.7 Flash67.7%estimated ± 7.2 pp, medium confidence
5Gemini 3.5 Flash67.7%estimated ± 7.2 pp, medium confidence
6Claude Opus 5.567.7%measured
7GPT-5.6 Sol67.7%estimated ± 7.2 pp, medium confidence
8Gemini 3.6 Flash67.7%estimated ± 7.2 pp, medium confidence
9GPT-6 Sol67.6%estimated ± 7.2 pp, medium confidence
10Qwen3.8 Max Preview67.6%estimated ± 7.2 pp, medium confidence
11Gemini 3.1 Pro67.4%estimated ± 7.2 pp, medium confidence
12Claude Opus 566.9%measured
13Claude Opus 4.866.2%measured
14Hy4 preview66.2%measured
15Claude Sonnet 5.565.6%measured
16Kimi K363.3%measured
17GLM-5.3-Flash62.4%measured
18GPT-5.6 Terra62.0%estimated ± 7.2 pp, medium confidence
19Muse Spark60.4%estimated ± 7.2 pp, medium confidence
20Qwen3.7 Plus60.4%estimated ± 7.2 pp, medium confidence
21Step 5 Preview60.3%measured
22Grok 4.559.6%estimated ± 7.2 pp, medium confidence
23Claude Fable 557.9%measured
24Gemini 3 Pro57.8%estimated ± 7.2 pp, medium confidence
25Qwen3.8-Flash-Next54.7%estimated ± 7.2 pp, medium confidence
26GPT-6 Luna54.1%estimated ± 7.2 pp, medium confidence
27GPT-5.554.1%measured
28GPT-5.453.2%measured
29Kimi K2.652.6%estimated ± 7.2 pp, medium confidence
30Apodex 1.151.9%estimated ± 7.2 pp, medium confidence
31Apodex 1.1 Mini51.9%estimated ± 7.2 pp, medium confidence
32Gemini 3.5 Flash-Lite51.4%estimated ± 7.2 pp, medium confidence
33Ling 3.0 Flash VL51.4%estimated ± 7.2 pp, medium confidence
34Gemini 3 Flash50.8%estimated ± 7.2 pp, medium confidence
35GPT-5.6 Luna50.8%estimated ± 7.2 pp, medium confidence
36GPT-5.3 Codex50.7%estimated ± 7.2 pp, medium confidence
37Grok 4.350.5%estimated ± 7.2 pp, medium confidence
38Qwen3.6 Plus50.5%estimated ± 7.2 pp, medium confidence
39Claude Sonnet 550.4%estimated ± 7.2 pp, medium confidence
40DeepSeek V4.1 Flash50.3%estimated ± 7.2 pp, medium confidence
41Claude Opus 4.750.3%estimated ± 7.2 pp, medium confidence
42Mistral Large 450.3%estimated ± 7.2 pp, medium confidence
43GPT-5.2-Codex50.3%estimated ± 7.2 pp, low confidence
44Qwen3.8-27B50.3%estimated ± 7.2 pp, low confidence
45GPT-5.150.3%estimated ± 7.2 pp, low confidence
46Claude Opus 4.6 (Adaptive)50.3%estimated ± 7.2 pp, low confidence
47Kimi K2.550.3%estimated ± 7.2 pp, low confidence
48Kimi K2.5 (Reasoning)50.3%estimated ± 7.2 pp, low confidence
49Step 3.7 Flash50.3%estimated ± 7.2 pp, low confidence
50Qwen3.5-122B-A10B50.3%estimated ± 7.2 pp, low confidence
51Qwen3.5-27B50.3%estimated ± 7.2 pp, low confidence
52Qwen3.6-35B-A3B50.3%estimated ± 7.2 pp, low confidence
53Gemini 2.5 Pro50.3%estimated ± 7.2 pp, low confidence
54Qwen3.6-27B50.3%estimated ± 7.2 pp, low confidence
55Claude Opus 4.5 Thinking50.3%estimated ± 7.2 pp, low confidence
56GPT-5 (high)50.3%estimated ± 7.2 pp, low confidence
57GPT-5 (medium)50.3%estimated ± 7.2 pp, low confidence
58Inkling-Small50.3%estimated ± 7.2 pp, low confidence
59Muse Glimmer 30B50.3%estimated ± 7.2 pp, low confidence
60Claude 3 Haiku50.3%estimated ± 7.2 pp, low confidence
61Claude 4.1 Opus Thinking50.3%estimated ± 7.2 pp, low confidence
62Claude 4 Sonnet50.3%estimated ± 7.2 pp, low confidence
63Claude Opus 4.550.3%estimated ± 7.2 pp, low confidence
64Claude Opus 4.650.3%estimated ± 7.2 pp, low confidence
65Claude Sonnet 4.650.3%estimated ± 7.2 pp, low confidence
66Command A+50.3%estimated ± 7.2 pp, low confidence
67Gemini 1.5 Pro50.3%estimated ± 7.2 pp, low confidence
68Gemini 2.5 Flash50.3%estimated ± 7.2 pp, low confidence
69Gemma 3 27B50.3%estimated ± 7.2 pp, low confidence
70Gemma 4 12B50.3%estimated ± 7.2 pp, low confidence
71Gemma 4 26B A4B50.3%estimated ± 7.2 pp, low confidence
72Gemma 4 31B50.3%estimated ± 7.2 pp, low confidence
73Gemma 4 E2B50.3%estimated ± 7.2 pp, low confidence
74Gemma 4 E4B50.3%estimated ± 7.2 pp, low confidence
75GLM-5V-Turbo50.3%estimated ± 7.2 pp, low confidence
76GPT-4.150.3%estimated ± 7.2 pp, low confidence
77GPT-4.1 mini50.3%estimated ± 7.2 pp, low confidence
78GPT-4.1 nano50.3%estimated ± 7.2 pp, low confidence
79GPT-4o mini50.3%estimated ± 7.2 pp, low confidence
80GPT-5.1-Codex50.3%estimated ± 7.2 pp, low confidence
81GPT-5.1-Codex-Max50.3%estimated ± 7.2 pp, low confidence
82GPT-5.4 mini50.3%estimated ± 7.2 pp, low confidence
83GPT-5.4 nano50.3%estimated ± 7.2 pp, low confidence
84Grok 450.3%estimated ± 7.2 pp, low confidence
85Grok 4.1 Fast50.3%estimated ± 7.2 pp, low confidence
86Grok 4.1 Fast (Reasoning)50.3%estimated ± 7.2 pp, low confidence
87Grok 4 Fast (Reasoning)50.3%estimated ± 7.2 pp, low confidence
88Inkling50.3%estimated ± 7.2 pp, low confidence
89LFM2.5-VL-1.6B-Extract50.3%estimated ± 7.2 pp, low confidence
90Llama 4 Maverick50.3%estimated ± 7.2 pp, low confidence
91Llama 4 Scout50.3%estimated ± 7.2 pp, low confidence
92MiMo-V2.6-Flash50.3%estimated ± 7.2 pp, low confidence
93MiMo-V2-Omni50.3%estimated ± 7.2 pp, low confidence
94Mistral Large 350.3%estimated ± 7.2 pp, low confidence
95Mistral Medium 350.3%estimated ± 7.2 pp, low confidence
96Mistral Medium 3.5 128B50.3%estimated ± 7.2 pp, low confidence
97Mistral Small 450.3%estimated ± 7.2 pp, low confidence
98Mistral Small 4 (Reasoning)50.3%estimated ± 7.2 pp, low confidence
99Nemotron 3 Nano Omni 30B A3B50.3%estimated ± 7.2 pp, low confidence
100Nova Pro50.3%estimated ± 7.2 pp, low confidence
101o350.3%estimated ± 7.2 pp, low confidence
102Phi-4 Multimodal Instruct50.3%estimated ± 7.2 pp, low confidence
103Qwen3.5-35B-A3B50.3%estimated ± 7.2 pp, low confidence
104Qwen3.5 397B50.3%estimated ± 7.2 pp, low confidence
105Qwen3.5 397B (Reasoning)50.3%estimated ± 7.2 pp, low confidence
106Qwen3-Omni-30B-A3B-Instruct50.3%estimated ± 7.2 pp, low confidence
107Qwen3-Omni-30B-A3B-Thinking50.3%estimated ± 7.2 pp, low confidence
108MiniMax M345.1%measured
109Claude Opus 4.7 (Adaptive)43.6%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General