benchgap
Agentic · tools

OSWorld 2.0 leaderboard

As of 2026-10-07, the highest measured score on OSWorld 2.0 is 72.6% by GPT-6 Astra. 99 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra72.6%measured
2Claude Opus 570.6%measured
3Gemini 4 Argon69.2%measured
4Atria Dawn Preview68.6%estimated ± 13.8 pp, low confidence
5GPT-5.5 Pro68.5%estimated ± 13.8 pp, low confidence
6GPT-5.4 Pro67.8%estimated ± 13.8 pp, low confidence
7Muse Spark 1.366.9%measured
8GPT-5.6 Sol62.6%measured
9DeepSeek V4.1 Flash61.0%estimated ± 13.3 pp, low confidence
10GPT-6 Sol60.5%measured
11Grok 4.659.6%estimated ± 13.3 pp, low confidence
12Gemini 3.8 Flash59.0%measured
13GLM-5.3-Flash59.0%estimated ± 5.5 pp, low confidence
14Grok 4.758.9%estimated ± 13.3 pp, low confidence
15GPT-6.1 Sol58.5%estimated ± 13.3 pp, low confidence
16Mistral Large 455.2%estimated ± 13.3 pp, low confidence
17GPT-6 Luna50.6%estimated ± 13.3 pp, low confidence
18GPT-5.6 Terra50.2%measured
19Claude Opus 5.548.7%measured
20Gemini 3.7 Flash47.9%measured
21GPT-5.6 Luna45.6%measured
22Claude Sonnet 5.545.2%estimated ± 5.5 pp, low confidence
23Claude Fable 5.141.7%measured
24Claude Mythos Preview40.2%estimated ± 13.6 pp, low confidence
25Claude Haiku 5.536.8%estimated ± 13.3 pp, low confidence
26Claude Opus 4.820.6%measured
27Claude Fable 520.4%estimated ± 3.6 pp, high confidence
28Claude Mythos 520.4%estimated ± 3.6 pp, high confidence
29Qwen3.8-27B19.6%estimated ± 3.6 pp, high confidence
30DeepSeek V4 Pro 081319.4%estimated ± 5.5 pp, low confidence
31Hy4 preview19.4%estimated ± 5.5 pp, low confidence
32Step 5 Preview19.4%estimated ± 5.5 pp, low confidence
33Qwen3.8-Flash-Next19.4%measured
34Qwen3.8 Max19.4%measured
35Kimi K319.4%estimated ± 5.5 pp, low confidence
36GLM-5.319.4%estimated ± 5.5 pp, low confidence
37Ornith-1.5-397B19.4%estimated ± 5.5 pp, low confidence
38DeepSeek V4 Flash 073119.4%estimated ± 5.5 pp, low confidence
39dots3-note Preview19.4%estimated ± 5.5 pp, low confidence
40Inkling-Small19.4%estimated ± 5.5 pp, low confidence
41Laguna S 2.119.4%estimated ± 5.5 pp, low confidence
42Ornith-1.5-35B-A3B19.4%estimated ± 5.5 pp, low confidence
43Ornith-1.5-9B19.4%estimated ± 5.5 pp, low confidence
44Claude Opus 4.7 (Adaptive)18.2%measured
45Gemini 3.6 Flash18.1%estimated ± 3.6 pp, high confidence
46Agents-A117.7%estimated ± 13.8 pp, low confidence
47Agents-A1-4B17.7%estimated ± 13.8 pp, low confidence
48Ling 3.0 Flash17.7%estimated ± 13.8 pp, low confidence
49LongCat-Flash-Lite-Sparse17.7%estimated ± 13.8 pp, low confidence
50Nemotron 3.5 Lightning 30B A3B NVFP417.7%estimated ± 13.8 pp, low confidence
51Beam17.7%estimated ± 13.8 pp, low confidence
52Solar Pro 417.7%estimated ± 13.8 pp, low confidence
53Holo3-35B-A3B17.7%estimated ± 3.6 pp, high confidence
54MiMo-V2.6-Pro17.0%estimated ± 3.6 pp, high confidence
55Claude Sonnet 516.1%estimated ± 3.6 pp, high confidence
56MiMo-V2.6-Flash15.7%estimated ± 3.6 pp, high confidence
57Muse Spark 1.114.2%measured
58Claude Opus 4.713.9%measured
59GLM-5.213.6%estimated ± 5.1 pp, low confidence
60Holo3-122B-A10B13.5%estimated ± 3.6 pp, high confidence
61GPT-5.513.0%measured
62Gemini 3.5 Flash13.0%estimated ± 3.6 pp, high confidence
63UI-Mate-27B11.4%estimated ± 3.6 pp, high confidence
64Qwen3.7 Max10.6%estimated ± 4.5 pp, medium confidence
65Gemini 3 Pro9.8%estimated ± 4.5 pp, medium confidence
66MiMo-V2.5-Pro9.4%estimated ± 4.5 pp, medium confidence
67GPT-5.49.2%estimated ± 3.6 pp, high confidence
68Claude Sonnet 4.68.3%measured
69Gemini 3.5 Flash-Lite8.1%estimated ± 3.6 pp, high confidence
70GLM-5.17.4%estimated ± 4.5 pp, medium confidence
71Grok 4.17.2%estimated ± 5.1 pp, low confidence
72Claude Opus 4.66.6%estimated ± 3.6 pp, high confidence
73Inkling6.2%estimated ± 13.3 pp, low confidence
74GPT-5.4 mini5.9%estimated ± 3.6 pp, high confidence
75Gemini 3.1 Pro4.9%estimated ± 4.5 pp, medium confidence
76Gemini 3 Flash4.7%estimated ± 4.5 pp, low confidence
77Kimi K2.64.6%measured
78MiniMax M34.6%measured
79Nemotron 3 Ultra3.7%estimated ± 13.3 pp, low confidence
80Qwen3.6-27B3.4%estimated ± 4.5 pp, low confidence
81Qwen3.7 Plus2.8%measured
82GPT-5.2-Codex1.0%estimated ± 4.5 pp, low confidence
83Step 3.7 Flash0.9%estimated ± 4.5 pp, low confidence
84GLM-50.4%estimated ± 4.5 pp, low confidence
85Qwen3.6 Plus0.1%estimated ± 4.5 pp, low confidence
86Claude 4 Sonnet0.0%estimated ± 4.5 pp, low confidence
87Claude Opus 4.50.0%estimated ± 3.6 pp, medium confidence
88Claude Sonnet 4.50.0%estimated ± 3.6 pp, medium confidence
89DeepSeek V3.20.0%estimated ± 4.5 pp, low confidence
90Gemini 2.5 Pro0.0%estimated ± 4.5 pp, low confidence
91Gemini 3.1 Flash-Lite0.0%estimated ± 4.5 pp, low confidence
92Gemma 4 31B0.0%estimated ± 4.5 pp, low confidence
93GLM-4.70.0%estimated ± 4.5 pp, low confidence
94GLM-5V-Turbo0.0%estimated ± 4.5 pp, low confidence
95GPT-4.10.0%estimated ± 4.5 pp, low confidence
96GPT-5.10.0%estimated ± 4.5 pp, low confidence
97GPT-5.1-Codex0.0%estimated ± 4.5 pp, low confidence
98GPT-5.20.0%estimated ± 3.6 pp, medium confidence
99GPT-5.3 Codex0.0%estimated ± 3.6 pp, medium confidence
100GPT-5.4 nano0.0%estimated ± 3.6 pp, medium confidence
101GPT-OSS 120B0.0%estimated ± 4.5 pp, low confidence
102Grok 40.0%estimated ± 4.5 pp, low confidence
103Grok 4.1 Fast0.0%estimated ± 4.5 pp, low confidence
104Grok 4.200.0%estimated ± 4.5 pp, low confidence
105Grok 4.30.0%estimated ± 4.5 pp, low confidence
106Grok Build 0.10.0%estimated ± 4.5 pp, low confidence
107Hy3 Preview0.0%estimated ± 4.5 pp, low confidence
108Kimi K2.50.0%estimated ± 4.5 pp, low confidence
109Kimi K2.5 (Reasoning)0.0%estimated ± 4.5 pp, low confidence
110MiMo-V2.50.0%estimated ± 4.5 pp, low confidence
111MiMo-V2-Pro0.0%estimated ± 4.5 pp, low confidence
112MiniMax M2.70.0%estimated ± 4.5 pp, low confidence
113Mistral Medium 3.5 128B0.0%estimated ± 4.5 pp, low confidence
114Muse Glimmer 30B0.0%estimated ± 3.6 pp, medium confidence
115Qwen3.5-122B-A10B0.0%estimated ± 3.6 pp, medium confidence
116Qwen3.5-27B0.0%estimated ± 3.6 pp, medium confidence
117Qwen3.5-35B-A3B0.0%estimated ± 3.6 pp, medium confidence
118Qwen3.5 397B0.0%estimated ± 4.5 pp, low confidence
119Qwen3.6-35B-A3B0.0%estimated ± 4.5 pp, low confidence
120Qwen3 Max0.0%estimated ± 4.5 pp, low confidence
121Trinity-Large-Thinking0.0%estimated ± 4.5 pp, low confidence
122UI-Mate-9B0.0%estimated ± 3.6 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General