benchgap
Agentic · tools

OSWorld-Verified leaderboard

As of 2026-10-07, the highest measured score on OSWorld-Verified is 86.1% by Qwen3.8 Max. 143 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.1100.0%estimated ± 1.9 pp, low confidence
2Claude Opus 5.5100.0%estimated ± 3.0 pp, medium confidence
3Gemini 4 Argon100.0%estimated ± 3.0 pp, medium confidence
4GPT-6 Sol100.0%estimated ± 3.0 pp, medium confidence
5GPT-6 Astra99.7%estimated ± 1.9 pp, low confidence
6Claude Sonnet 5.590.7%estimated ± 8.7 pp, low confidence
7Grok 4.787.5%estimated ± 8.7 pp, low confidence
8Claude Opus 586.4%estimated ± 1.9 pp, low confidence
9Qwen3.8 Max Preview86.3%estimated ± 8.7 pp, low confidence
10Qwen3.8 Max86.1%measured
11GPT-5.5 Pro85.3%estimated ± 5.4 pp, low confidence
12Claude Fable 585.0%measured
13Claude Mythos 585.0%measured
14GPT-5.4 Pro84.3%estimated ± 5.4 pp, low confidence
15Qwen3.8-27B84.3%measured
16GPT-6.1 Sol83.8%estimated ± 8.7 pp, low confidence
17Claude Opus 4.883.4%measured
18Gemini 3.6 Flash83.0%measured
19Qwen3.8-Flash-Next82.8%estimated ± 3.0 pp, high confidence
20Holo3-35B-A3B82.6%measured
21GPT-5.6 Sol82.2%estimated ± 1.9 pp, medium confidence
22MiMo-V2.6-Pro82.0%measured
23Gemini 3.8 Flash81.3%estimated ± 1.9 pp, medium confidence
24Muse Spark 1.281.2%estimated ± 8.7 pp, low confidence
25Claude Sonnet 581.2%measured
26Ornith-1.5-397B80.9%estimated ± 5.4 pp, medium confidence
27MiMo-V2.6-Flash80.8%measured
28Muse Spark 1.180.8%measured
29DeepSeek V4.1 Flash80.7%estimated ± 4.7 pp, high confidence
30Ling 3.1 Flash80.7%estimated ± 4.7 pp, high confidence
31Fugu Cyber80.5%estimated ± 4.7 pp, high confidence
32Atria Dawn Preview80.4%estimated ± 4.7 pp, high confidence
33Gemini 3.8 Flash Cyber80.3%estimated ± 4.7 pp, high confidence
34GPT-6 Luna80.2%estimated ± 8.7 pp, low confidence
35Step 5 Preview80.0%estimated ± 4.7 pp, high confidence
36GLM-5.380.0%estimated ± 4.7 pp, high confidence
37Mistral Large 479.8%estimated ± 8.7 pp, low confidence
38DeepSeek V4 Pro 081379.7%estimated ± 4.7 pp, high confidence
39Gemini 3.5 Flash Cyber79.7%estimated ± 4.7 pp, high confidence
40Claude Mythos Preview79.7%estimated ± 4.7 pp, high confidence
41Apodex 1.179.3%estimated ± 6.7 pp, low confidence
42Muse Spark 1.379.3%estimated ± 1.9 pp, medium confidence
43Grok 4.578.9%estimated ± 8.7 pp, low confidence
44Holo3-122B-A10B78.9%measured
45Kimi K378.8%estimated ± 1.9 pp, medium confidence
46GPT-5.578.7%measured
47Hy4 preview78.6%estimated ± 4.7 pp, high confidence
48GLM-5.278.5%estimated ± 8.7 pp, low confidence
49Gemini 3.5 Flash78.4%measured
50DeepSeek V4 Flash 073178.2%estimated ± 4.7 pp, high confidence
51Gemini 3.7 Flash78.0%estimated ± 1.9 pp, medium confidence
52GPT-5.6 Terra78.0%estimated ± 1.9 pp, medium confidence
53Claude Opus 4.7 (Adaptive)78.0%measured
54UI-Mate-27B77.0%measured
55Grok 4.676.8%estimated ± 1.9 pp, medium confidence
56dots3-note Preview76.6%estimated ± 5.4 pp, medium confidence
57GLM-5.176.0%estimated ± 4.7 pp, high confidence
58GPT-5.475.0%measured
59Claude Opus 4.774.2%estimated ± 1.9 pp, medium confidence
60GPT-5.6 Luna74.2%estimated ± 1.9 pp, medium confidence
61Qwen3.7 Max74.0%estimated ± 5.9 pp, medium confidence
62Gemini 3.5 Flash-Lite74.0%measured
63Apodex 1.1 Mini73.8%estimated ± 8.7 pp, low confidence
64Quasar 438B73.5%estimated ± 8.7 pp, low confidence
65Qwen3.7 Plus73.3%measured
66Kimi K2.673.1%measured
67Gemini 3 Pro73.0%estimated ± 5.9 pp, medium confidence
68Ling 3.0 Flash VL73.0%estimated ± 8.7 pp, low confidence
69Claude Opus 4.672.7%measured
70MiMo-V2.5-Pro72.5%estimated ± 5.9 pp, medium confidence
71Claude Sonnet 4.672.1%measured
72GPT-5.4 mini72.1%measured
73Hy370.4%estimated ± 8.7 pp, low confidence
74MiniMax M370.1%measured
75Kimi K2.7 Code69.7%estimated ± 8.7 pp, low confidence
76Inkling-Small69.1%estimated ± 5.4 pp, medium confidence
77Beam69.1%estimated ± 5.4 pp, medium confidence
78GLM-5.3-Flash69.0%estimated ± 6.2 pp, low confidence
79Inkling68.7%estimated ± 5.4 pp, medium confidence
80A.X K267.7%estimated ± 8.7 pp, low confidence
81Ling 3.0 Flash FP867.3%estimated ± 8.7 pp, low confidence
82Step 3.7 Flash67.1%estimated ± 5.4 pp, medium confidence
83Agents-A166.8%estimated ± 5.4 pp, medium confidence
84Gemini 3.1 Pro66.8%estimated ± 5.9 pp, medium confidence
85Gemini 3 Flash66.5%estimated ± 5.9 pp, medium confidence
86Claude Opus 4.566.3%measured
87UI-Mate-9B66.2%measured
88Solar Open 266.0%estimated ± 10.9 pp, low confidence
89Muse Glimmer 30B65.9%measured
90Muse Spark65.8%estimated ± 4.7 pp, medium confidence
91GLM-565.6%estimated ± 4.7 pp, medium confidence
92Qwen3.6-27B64.8%estimated ± 5.9 pp, medium confidence
93GPT-5.3 Codex64.7%measured
94Ling 3.0 Flash63.0%estimated ± 5.4 pp, medium confidence
95GPT-5.2-Codex61.9%estimated ± 5.9 pp, medium confidence
96Claude Sonnet 4.561.4%measured
97MiniCPM5-2B61.1%estimated ± 8.7 pp, low confidence
98Qwen3.6 Plus60.9%estimated ± 5.9 pp, medium confidence
99LLaDA2.2-flash60.2%estimated ± 10.9 pp, low confidence
100GPT-5.1-Codex60.1%estimated ± 5.9 pp, medium confidence
101Grok Build 0.159.7%estimated ± 5.9 pp, medium confidence
102MiMo-V2-Flash59.0%estimated ± 8.7 pp, low confidence
103Grok 4.1 Fast58.3%estimated ± 5.9 pp, medium confidence
104Ornith-1.5-35B-A3B58.3%estimated ± 5.4 pp, medium confidence
105MiMo-V2.558.0%estimated ± 5.9 pp, medium confidence
106Qwen3.5-122B-A10B58.0%measured
107Granite 4.2 30B57.6%estimated ± 8.7 pp, low confidence
108Agents-A1-4B57.6%estimated ± 5.4 pp, medium confidence
109Gemma 4 26B A4B57.3%estimated ± 8.7 pp, low confidence
110Ling 3.0 Tiny57.2%estimated ± 8.7 pp, low confidence
111Qwen3.5-27B56.2%measured
112Grok 4.356.2%estimated ± 5.9 pp, medium confidence
113Qwen3 Max56.1%estimated ± 5.9 pp, medium confidence
114Command A+56.0%estimated ± 8.7 pp, low confidence
115Celeris-155.6%estimated ± 8.7 pp, low confidence
116DeepSeek V355.6%estimated ± 8.7 pp, low confidence
117DeepSeek V3 032455.6%estimated ± 8.7 pp, low confidence
118Gemma 3 27B55.6%estimated ± 8.7 pp, low confidence
119Gemma 4 12B55.6%estimated ± 8.7 pp, low confidence
120Gemma 4 E2B55.6%estimated ± 8.7 pp, low confidence
121Gemma 4 E4B55.6%estimated ± 8.7 pp, low confidence
122GPT-4.1 mini55.6%estimated ± 8.7 pp, low confidence
123GPT-4.1 nano55.6%estimated ± 8.7 pp, low confidence
124GPT-4o55.6%estimated ± 8.7 pp, low confidence
125GPT-4o mini55.6%estimated ± 8.7 pp, low confidence
126GPT-OSS 20B55.6%estimated ± 8.7 pp, low confidence
127Granite 4.2 3B55.6%estimated ± 8.7 pp, low confidence
128Granite 4.2 8B55.6%estimated ± 8.7 pp, low confidence
129K-Exaone55.6%estimated ± 8.7 pp, low confidence
130LFM2.5-2.6B55.6%estimated ± 8.7 pp, low confidence
131Ling 2.6 Flash55.6%estimated ± 8.7 pp, low confidence
132Llama 4 Maverick55.6%estimated ± 8.7 pp, low confidence
133Llama 4 Scout55.6%estimated ± 8.7 pp, low confidence
134Mistral Large 355.6%estimated ± 8.7 pp, low confidence
135Mistral Small 455.6%estimated ± 8.7 pp, low confidence
136Mistral Small 4 (Reasoning)55.6%estimated ± 8.7 pp, low confidence
137Nemotron 3 Nano 30B55.6%estimated ± 8.7 pp, low confidence
138Nemotron 3 Nano Omni 30B A3B55.6%estimated ± 8.7 pp, low confidence
139Nemotron 3 Super 100B55.6%estimated ± 8.7 pp, low confidence
140North Mini Code55.6%estimated ± 8.7 pp, low confidence
141Solar Pro 355.6%estimated ± 8.7 pp, low confidence
142Trinity-Large-Preview55.6%estimated ± 8.7 pp, low confidence
143Ultravox v0.6 Llama 3.3 70B55.6%estimated ± 8.7 pp, low confidence
144Qwen3.6-35B-A3B55.5%estimated ± 5.9 pp, medium confidence
145Claude 4.1 Opus55.4%estimated ± 8.7 pp, low confidence
146Grok 455.4%estimated ± 5.9 pp, medium confidence
147Gemini 2.5 Pro55.2%estimated ± 5.9 pp, medium confidence
148GPT-5.154.9%estimated ± 5.9 pp, medium confidence
149MiniMax M2.754.5%estimated ± 5.9 pp, medium confidence
150Qwen3.5-35B-A3B54.5%measured
151Claude 4 Sonnet54.2%estimated ± 5.9 pp, medium confidence
152Mistral Medium 3.5 128B54.0%estimated ± 5.9 pp, medium confidence
153Gemini 3.1 Flash-Lite53.8%estimated ± 5.9 pp, medium confidence
154Grok 4.2053.8%estimated ± 5.9 pp, medium confidence
155Qwen3.5 397B53.8%estimated ± 5.4 pp, medium confidence
156Hy3 Preview53.4%estimated ± 5.9 pp, medium confidence
157MiMo-V2-Pro53.3%estimated ± 5.9 pp, medium confidence
158Gemma 4 31B53.0%estimated ± 5.9 pp, medium confidence
159Kimi K2.552.8%estimated ± 5.4 pp, low confidence
160Kimi K2.5 (Reasoning)52.8%estimated ± 5.4 pp, low confidence
161Trinity-Large-Thinking52.5%estimated ± 5.9 pp, medium confidence
162GLM-5V-Turbo52.2%estimated ± 5.9 pp, medium confidence
163GPT-OSS 120B52.1%estimated ± 5.9 pp, medium confidence
164DeepSeek V3.252.1%estimated ± 5.9 pp, medium confidence
165GPT-4.151.8%estimated ± 5.9 pp, low confidence
166Ornith-1.5-9B50.5%estimated ± 5.4 pp, low confidence
167Qwen3.5 Plus50.5%estimated ± 8.7 pp, low confidence
168GLM-4.748.8%estimated ± 5.4 pp, low confidence
169Solar Pro 447.9%estimated ± 5.4 pp, low confidence
170LongCat-Flash-Lite-Sparse47.8%estimated ± 5.4 pp, low confidence
171GPT-5.247.3%measured
172Nemotron 3 Ultra47.0%estimated ± 5.4 pp, low confidence
173Claude Haiku 4.546.3%estimated ± 8.7 pp, low confidence
174Mercury 2.546.2%estimated ± 6.7 pp, low confidence
175Nemotron 3.5 Lightning 30B A3B NVFP446.2%estimated ± 5.4 pp, low confidence
176GPT-5.4 nano39.0%measured
177GPT-5 (high)30.2%estimated ± 8.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General