benchgap
Vision & documents

ScreenSpot Pro leaderboard

As of 2026-10-07, the highest measured score on ScreenSpot Pro is 92.7% by GPT-6 Astra. 108 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra92.7%measured
2Claude Mythos 589.8%estimated ± 7.1 pp, medium confidence
3Claude Opus 4.887.9%measured
4Claude Opus 5.587.8%estimated ± 10.7 pp, low confidence
5Claude Opus 4.7 (Adaptive)87.3%estimated ± 7.1 pp, medium confidence
6GPT-6.1 Sol85.9%estimated ± 10.7 pp, low confidence
7GLM-5.3-Flash85.5%estimated ± 7.1 pp, medium confidence
8Gemini 3.8 Flash85.4%estimated ± 10.7 pp, low confidence
9GPT-5.485.4%measured
10GPT-5.6 Sol85.3%estimated ± 10.0 pp, low confidence
11LFM2.5-VL-3B85.1%estimated ± 2.5 pp, low confidence
12Qwen3.6-27B85.0%estimated ± 2.5 pp, low confidence
13Grok 4.2085.0%estimated ± 2.5 pp, low confidence
14Qwen3.6-35B-A3B85.0%estimated ± 2.5 pp, low confidence
15Gemini 3.7 Flash84.7%estimated ± 7.1 pp, medium confidence
16dots3-note Preview84.6%estimated ± 2.5 pp, medium confidence
17Qwen3.8 Max84.5%measured
18Claude Opus 584.4%estimated ± 10.7 pp, low confidence
19Gemini 3.1 Pro84.4%measured
20Muse Spark 1.184.3%estimated ± 7.1 pp, medium confidence
21Claude Sonnet 584.2%estimated ± 7.1 pp, medium confidence
22Muse Spark84.1%measured
23Claude Opus 4.683.1%measured
24Gemini 3.6 Flash82.7%estimated ± 10.7 pp, low confidence
25Kimi K382.4%estimated ± 4.5 pp, medium confidence
26GPT-6 Sol82.4%estimated ± 10.7 pp, low confidence
27Qwen3.8 Max Preview82.3%estimated ± 10.7 pp, low confidence
28Sakana Fugu-Ultra82.0%estimated ± 7.1 pp, medium confidence
29Step 3.7 Flash81.9%estimated ± 2.5 pp, medium confidence
30GPT-5.580.8%estimated ± 10.0 pp, low confidence
31Sakana Fugu80.0%estimated ± 7.1 pp, medium confidence
32Grok 4.579.6%estimated ± 10.7 pp, low confidence
33GPT-5.6 Terra79.5%estimated ± 10.0 pp, low confidence
34Seed 2.1 Pro79.3%estimated ± 4.5 pp, medium confidence
35Qwen3.7 Plus79.0%measured
36GPT-6 Luna78.8%estimated ± 10.7 pp, low confidence
37Gemini 3.5 Flash78.7%estimated ± 7.1 pp, medium confidence
38Apodex 1.178.2%estimated ± 10.7 pp, low confidence
39Apodex 1.1 Mini78.2%estimated ± 10.7 pp, low confidence
40Gemini 3.5 Flash-Lite78.0%estimated ± 10.7 pp, low confidence
41Ling 3.0 Flash VL78.0%estimated ± 10.7 pp, low confidence
42Qwen3.8-Omni-Flash77.8%estimated ± 4.5 pp, medium confidence
43Gemini 3 Flash77.6%estimated ± 10.7 pp, low confidence
44GPT-5.3 Codex77.5%estimated ± 10.7 pp, low confidence
45DeepSeek V4.1 Flash75.8%estimated ± 10.7 pp, low confidence
46Qwen3.8-Flash-Next75.6%estimated ± 4.5 pp, medium confidence
47Muse Glimmer 30B75.4%measured
48Inkling75.2%estimated ± 7.1 pp, medium confidence
49Claude Opus 4.775.2%estimated ± 10.7 pp, low confidence
50Mistral Large 475.2%estimated ± 10.7 pp, low confidence
51GPT-5.2-Codex75.1%estimated ± 10.7 pp, low confidence
52Seed 2.1 Turbo74.7%estimated ± 4.5 pp, medium confidence
53Qwen3.8-27B74.5%estimated ± 4.5 pp, medium confidence
54GPT-5.174.2%estimated ± 10.7 pp, low confidence
55Kimi K2.574.2%estimated ± 10.0 pp, low confidence
56Kimi K2.5 (Reasoning)74.2%estimated ± 10.0 pp, low confidence
57Claude Opus 4.6 (Adaptive)74.1%estimated ± 10.7 pp, low confidence
58Inkling-Small74.1%estimated ± 7.1 pp, medium confidence
59GPT-5.6 Luna74.0%estimated ± 10.0 pp, low confidence
60Gemini 2.5 Pro73.6%estimated ± 10.7 pp, low confidence
61MiMo-V2.573.6%estimated ± 7.1 pp, medium confidence
62Grok 4.373.3%estimated ± 10.0 pp, low confidence
63MiniMax M373.3%estimated ± 10.0 pp, low confidence
64Pareto 26.973.0%estimated ± 10.0 pp, low confidence
65GPT-5 (medium)73.0%estimated ± 10.7 pp, low confidence
66GPT-5 (high)72.9%estimated ± 10.7 pp, low confidence
67Gemini 3 Pro72.7%measured
68Claude Opus 4.5 Thinking72.7%estimated ± 10.7 pp, low confidence
69MiMo-V2.6-Flash71.8%estimated ± 10.7 pp, low confidence
70GLM-5V-Turbo71.5%estimated ± 10.7 pp, low confidence
71GPT-5.1-Codex71.2%estimated ± 10.7 pp, low confidence
72GPT-5.1-Codex-Max71.2%estimated ± 10.7 pp, low confidence
73Holo2-235B-A22B70.6%measured
74Gemma 4 31B70.4%estimated ± 10.0 pp, low confidence
75GPT-5.4 mini69.7%estimated ± 10.0 pp, low confidence
76Kimi K2.669.7%estimated ± 4.5 pp, medium confidence
77o368.9%estimated ± 10.7 pp, low confidence
78MiMo-V2-Omni68.7%estimated ± 10.7 pp, low confidence
79Step 5 Preview68.3%estimated ± 10.0 pp, low confidence
80Qwen3.6 Plus68.2%measured
81Grok 467.8%estimated ± 10.7 pp, low confidence
82Qwen3.5-122B-A10B67.5%estimated ± 4.5 pp, medium confidence
83Qwen3.5-27B67.1%estimated ± 4.5 pp, medium confidence
84Claude 4.1 Opus Thinking67.0%estimated ± 10.7 pp, low confidence
85Claude Sonnet 4.667.0%estimated ± 7.1 pp, medium confidence
86Holo2-30B-A3B66.1%measured
87Qwen3.5 397B65.6%measured
88Gemini 2.5 Flash65.1%estimated ± 10.7 pp, low confidence
89Mistral Medium 3.5 128B64.7%estimated ± 10.7 pp, low confidence
90Grok 4.1 Fast (Reasoning)63.6%estimated ± 10.7 pp, low confidence
91Gemma 4 26B A4B63.3%estimated ± 10.0 pp, low confidence
92Qwen3.5-35B-A3B63.3%estimated ± 4.5 pp, medium confidence
93Claude 4 Sonnet63.0%estimated ± 10.7 pp, low confidence
94Llama 4 Maverick62.8%estimated ± 10.7 pp, low confidence
95Grok 4 Fast (Reasoning)62.6%estimated ± 10.7 pp, low confidence
96GPT-4.162.3%estimated ± 10.7 pp, low confidence
97Qwen3-Omni-30B-A3B-Thinking61.7%estimated ± 10.7 pp, low confidence
98GPT-5.261.6%estimated ± 4.5 pp, medium confidence
99GPT-4.1 mini60.9%estimated ± 10.7 pp, low confidence
100Mistral Small 460.1%estimated ± 10.7 pp, low confidence
101Mistral Small 4 (Reasoning)60.1%estimated ± 10.7 pp, low confidence
102Mistral Large 359.6%estimated ± 10.7 pp, low confidence
103Qwen3-Omni-30B-A3B-Instruct59.5%estimated ± 10.7 pp, low confidence
104Gemini 1.5 Pro59.3%estimated ± 10.7 pp, low confidence
105Holo2-8B58.9%measured
106Mistral Medium 358.6%estimated ± 10.7 pp, low confidence
107Llama 4 Scout58.6%estimated ± 10.7 pp, low confidence
108Qwen3.5 397B (Reasoning)58.5%estimated ± 10.7 pp, low confidence
109Gemini 3.1 Flash-Lite58.3%estimated ± 7.1 pp, medium confidence
110Gemma 4 E4B58.1%estimated ± 10.7 pp, low confidence
111Nemotron 3 Nano Omni 30B A3B57.8%measured
112Grok 4.1 Fast57.4%estimated ± 10.7 pp, low confidence
113Gemma 3 27B57.3%estimated ± 10.7 pp, low confidence
114Interfaze Beta57.3%estimated ± 10.0 pp, low confidence
115Holo2-4B57.2%measured
116Gemma 4 E2B56.7%estimated ± 10.7 pp, low confidence
117Nova Pro56.7%estimated ± 10.7 pp, low confidence
118GPT-4o mini56.4%estimated ± 10.7 pp, low confidence
119GPT-4.1 nano56.2%estimated ± 10.7 pp, low confidence
120Claude 3 Haiku55.8%estimated ± 10.7 pp, low confidence
121LFM2.5-VL-1.6B-Extract55.7%estimated ± 10.7 pp, low confidence
122Phi-4 Multimodal Instruct55.6%estimated ± 10.7 pp, low confidence
123Gemma 4 12B55.5%estimated ± 4.5 pp, medium confidence
124GPT-5.4 nano46.7%estimated ± 10.0 pp, low confidence
125Claude Opus 4.545.7%measured
126Command A+15.3%estimated ± 7.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General