benchgap
Agentic · tools

BrowseComp leaderboard

As of 2026-10-07, the highest measured score on BrowseComp is 92.5% by Atria Dawn Preview. 143 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Atria Dawn Preview92.5%measured
2GPT-5.6 Sol92.2%measured
3GPT-6 Astra91.5%measured
4Gemini 4 Argon91.5%estimated ± 2.2 pp, high confidence
5Claude Fable 591.2%estimated ± 2.8 pp, high confidence
6Kimi K391.2%measured
7Muse Spark 1.391.1%estimated ± 2.2 pp, high confidence
8Claude Opus 590.8%measured
9GPT-6.1 Sol90.7%estimated ± 1.8 pp, medium confidence
10GPT-5.5 Pro90.1%measured
11Gemini 3.8 Flash89.4%estimated ± 2.2 pp, high confidence
12GPT-5.4 Pro89.3%measured
13GLM-5.3-Flash89.0%estimated ± 4.7 pp, low confidence
14Ling 3.1 Flash88.9%estimated ± 5.2 pp, low confidence
15Step 5 Preview88.7%measured
16Fugu Cyber88.3%estimated ± 5.2 pp, low confidence
17Claude Mythos 588.0%measured
18Gemini 3.8 Flash Cyber87.9%estimated ± 5.2 pp, low confidence
19GPT-5.6 Terra87.5%measured
20GPT-6 Sol87.0%estimated ± 1.8 pp, medium confidence
21Claude Sonnet 5.587.0%estimated ± 6.4 pp, medium confidence
22Ornith-1.5-397B86.6%measured
23Claude Opus 5.586.3%estimated ± 2.2 pp, high confidence
24Gemini 3.5 Flash Cyber86.2%estimated ± 5.2 pp, low confidence
25Qwen3.8 Max Preview86.2%estimated ± 3.4 pp, medium confidence
26Gemini 3.7 Flash86.1%estimated ± 2.2 pp, high confidence
27GLM-5.285.8%estimated ± 3.4 pp, high confidence
28MiMo-V2.6-Pro85.8%estimated ± 1.8 pp, medium confidence
29Claude Opus 4.6 (Adaptive)85.7%estimated ± 3.4 pp, high confidence
30Claude Mythos Preview85.7%estimated ± 1.8 pp, medium confidence
31Gemini 3.1 Pro85.6%estimated ± 3.4 pp, high confidence
32Apodex 1.185.6%estimated ± 3.4 pp, high confidence
33Apodex 1.1 Mini85.6%estimated ± 3.4 pp, high confidence
34DeepSeek V4.1 Flash85.1%estimated ± 1.8 pp, medium confidence
35GLM-5.385.0%estimated ± 1.8 pp, medium confidence
36Muse Spark 1.284.9%estimated ± 9.0 pp, medium confidence
37Claude Sonnet 584.7%measured
38Claude Haiku 5.584.6%estimated ± 7.6 pp, medium confidence
39Claude Fable 5.184.5%estimated ± 2.2 pp, high confidence
40Grok 4.684.5%estimated ± 2.8 pp, high confidence
41GPT-5.584.4%measured
42Claude Opus 4.884.3%measured
43Grok 4.584.2%estimated ± 9.0 pp, medium confidence
44GPT-6 Luna84.0%estimated ± 1.8 pp, medium confidence
45Gemini 3.6 Flash83.8%estimated ± 3.1 pp, high confidence
46Holo3-35B-A3B83.8%estimated ± 3.1 pp, high confidence
47Qwen3.8-27B83.8%estimated ± 3.1 pp, high confidence
48Holo3-122B-A10B83.8%estimated ± 3.1 pp, high confidence
49Gemini 3.5 Flash83.8%estimated ± 3.1 pp, high confidence
50UI-Mate-27B83.8%estimated ± 3.1 pp, high confidence
51Gemini 3.5 Flash-Lite83.8%estimated ± 3.1 pp, high confidence
52GPT-5.4 mini83.7%estimated ± 3.1 pp, high confidence
53Claude Opus 4.683.7%measured
54MiniMax M383.5%measured
55DeepSeek V4 Pro 081383.4%measured
56Hy4 preview83.3%estimated ± 5.2 pp, low confidence
57dots3-note Preview83.3%measured
58GPT-5.6 Luna83.3%measured
59Kimi K2.683.2%measured
60Qwen3.8-Flash-Next82.8%estimated ± 2.2 pp, high confidence
61Qwen3.8 Max82.8%estimated ± 2.2 pp, high confidence
62Claude Opus 4.782.8%estimated ± 2.2 pp, high confidence
63Claude Sonnet 4.682.8%estimated ± 2.2 pp, high confidence
64Qwen3.7 Plus82.8%estimated ± 2.2 pp, medium confidence
65Mistral Large 482.7%estimated ± 10.1 pp, low confidence
66GPT-5.482.7%measured
67MiMo-V2.6-Flash82.4%estimated ± 1.8 pp, medium confidence
68K-EXAONE 2.082.0%estimated ± 9.4 pp, low confidence
69Ornith-1.0-397B81.6%estimated ± 9.4 pp, low confidence
70Qwen3.7 Max81.1%estimated ± 5.7 pp, medium confidence
71Gemini 3 Pro81.1%estimated ± 5.7 pp, medium confidence
72Muse Spark 1.180.9%estimated ± 1.8 pp, low confidence
73Gemini 3 Flash80.7%estimated ± 5.7 pp, medium confidence
74Qwen3.6-27B80.0%estimated ± 5.7 pp, medium confidence
75Quasar 438B79.8%estimated ± 9.0 pp, medium confidence
76Kimi K2.7 Code79.8%estimated ± 8.4 pp, medium confidence
77Claude Opus 4.579.4%estimated ± 3.1 pp, high confidence
78Claude Opus 4.7 (Adaptive)79.3%measured
79Grok 4.779.1%estimated ± 4.7 pp, low confidence
80UI-Mate-9B79.1%estimated ± 3.1 pp, high confidence
81Grok 4.378.7%estimated ± 3.4 pp, high confidence
82Muse Glimmer 30B78.3%estimated ± 3.1 pp, high confidence
83Inkling-Small77.4%measured
84Beam77.4%measured
85Inkling77.1%measured
86Ornith-1.0-35B76.9%estimated ± 9.4 pp, low confidence
87GPT-5.2-Codex76.5%estimated ± 5.7 pp, medium confidence
88Mercury 2.576.2%estimated ± 8.2 pp, low confidence
89Step 3.7 Flash75.8%measured
90Agents-A175.5%measured
91Ling 3.0 Flash VL75.3%estimated ± 10.1 pp, low confidence
92Hy375.1%estimated ± 9.0 pp, medium confidence
93GPT-5.3 Codex74.3%estimated ± 3.1 pp, high confidence
94GLM-573.4%estimated ± 3.4 pp, high confidence
95DeepSeek V4 Flash 073173.2%measured
96Ling 3.0 Flash72.2%measured
97Ornith-1.0-9B72.2%estimated ± 9.4 pp, low confidence
98LFM2.5-2.6B72.0%estimated ± 9.4 pp, low confidence
99Laguna S 2.171.7%estimated ± 6.4 pp, medium confidence
100GPT-5.1-Codex71.2%estimated ± 5.7 pp, medium confidence
101Ling 3.0 Flash FP870.7%estimated ± 9.0 pp, medium confidence
102Grok Build 0.169.6%estimated ± 5.7 pp, medium confidence
103Claude 4.1 Opus68.4%estimated ± 6.2 pp, medium confidence
104GLM-5.168.0%measured
105LLaDA2.2-mini67.7%estimated ± 9.4 pp, low confidence
106Ornith-1.5-35B-A3B67.6%measured
107Qwen3.6 Plus66.9%estimated ± 4.4 pp, high confidence
108Agents-A1-4B66.8%measured
109GLM-5-Turbo66.6%estimated ± 9.4 pp, low confidence
110A.X K266.2%estimated ± 10.1 pp, low confidence
111Qwen3.5 Plus66.2%estimated ± 6.2 pp, medium confidence
112GPT-5.265.8%measured
113Claude Sonnet 4.565.1%estimated ± 3.1 pp, high confidence
114Grok 4.1 Fast64.6%estimated ± 5.7 pp, medium confidence
115Claude Haiku 4.564.5%estimated ± 6.2 pp, medium confidence
116Solar Open 263.9%estimated ± 8.4 pp, medium confidence
117Qwen3.5-122B-A10B63.8%measured
118MiMo-V2.563.6%estimated ± 5.7 pp, medium confidence
119GPT-5.4 nano62.8%estimated ± 3.1 pp, medium confidence
120Qwen3.5 397B62.0%measured
121Qwen3.5-27B61.0%measured
122Qwen3.5-35B-A3B61.0%measured
123Kimi K2.560.6%measured
124Kimi K2.5 (Reasoning)60.6%measured
125Qwen3 Max59.8%estimated ± 5.7 pp, medium confidence
126GPT-5 (high)59.6%estimated ± 6.2 pp, low confidence
127Grok 459.2%estimated ± 5.7 pp, medium confidence
128Gemini 2.5 Pro59.1%estimated ± 5.7 pp, medium confidence
129LLaDA2.2-flash59.0%estimated ± 8.4 pp, medium confidence
130GPT-5.158.9%estimated ± 5.7 pp, medium confidence
131Qwen3.6-35B-A3B58.9%estimated ± 4.4 pp, medium confidence
132Claude 4 Sonnet58.8%estimated ± 5.7 pp, medium confidence
133Mistral Medium 3.5 128B58.7%estimated ± 5.7 pp, medium confidence
134Gemini 3.1 Flash-Lite58.7%estimated ± 5.7 pp, medium confidence
135Grok 4.2058.7%estimated ± 5.7 pp, medium confidence
136Hy3 Preview58.7%estimated ± 5.7 pp, medium confidence
137MiMo-V2-Pro58.7%estimated ± 5.7 pp, medium confidence
138Gemma 4 31B58.6%estimated ± 5.7 pp, medium confidence
139Trinity-Large-Thinking58.6%estimated ± 5.7 pp, medium confidence
140GLM-5V-Turbo58.6%estimated ± 5.7 pp, medium confidence
141DeepSeek V3.258.6%estimated ± 5.7 pp, medium confidence
142GPT-4.158.6%estimated ± 5.7 pp, low confidence
143MiMo-V2-Omni57.6%estimated ± 9.4 pp, low confidence
144Muse Spark56.7%estimated ± 5.2 pp, low confidence
145Ornith-1.5-9B56.4%measured
146MiniMax M2.755.2%estimated ± 3.4 pp, medium confidence
147GLM-4.752.0%measured
148Solar Pro 449.2%measured
149LongCat-Flash-Lite-Sparse48.6%measured
150Grok 4.146.0%estimated ± 7.1 pp, low confidence
151MiniCPM5-2B44.6%estimated ± 10.1 pp, low confidence
152Nemotron 3 Ultra44.4%measured
153Nemotron 3.5 Lightning 30B A3B NVFP436.8%measured
154MiMo-V2-Flash32.9%estimated ± 10.1 pp, low confidence
155Granite 4.2 8B28.5%estimated ± 9.0 pp, low confidence
156Command A+28.0%estimated ± 9.0 pp, low confidence
157Granite 4.2 30B22.2%estimated ± 10.1 pp, low confidence
158Mistral Large 320.7%estimated ± 9.0 pp, low confidence
159Gemma 4 26B A4B19.8%estimated ± 10.1 pp, low confidence
160Ling 3.0 Tiny18.9%estimated ± 10.1 pp, low confidence
161Mistral Small 413.3%estimated ± 9.0 pp, low confidence
162Mistral Small 4 (Reasoning)13.3%estimated ± 9.0 pp, low confidence
163Trinity-Large-Preview11.0%estimated ± 9.0 pp, low confidence
164Nemotron 3 Nano 30B9.6%estimated ± 9.0 pp, low confidence
165DeepSeek V37.7%estimated ± 9.0 pp, low confidence
166Celeris-16.4%estimated ± 9.0 pp, low confidence
167Llama 4 Maverick6.2%estimated ± 9.0 pp, low confidence
168Llama 4 Scout5.5%estimated ± 9.0 pp, low confidence
169GPT-OSS 120B1.6%estimated ± 3.4 pp, medium confidence
170Gemma 3 27B1.5%estimated ± 9.0 pp, low confidence
171MiMo-V2.5-Pro0.6%estimated ± 3.4 pp, medium confidence
172Nemotron 3 Super 100B0.2%estimated ± 3.4 pp, medium confidence
173GPT-OSS 20B0.0%estimated ± 3.4 pp, medium confidence
174DeepSeek V3 03240.0%estimated ± 10.1 pp, low confidence
175Gemma 4 12B0.0%estimated ± 10.1 pp, low confidence
176Gemma 4 E2B0.0%estimated ± 10.1 pp, low confidence
177Gemma 4 E4B0.0%estimated ± 10.1 pp, low confidence
178GPT-4.1 mini0.0%estimated ± 10.1 pp, low confidence
179GPT-4.1 nano0.0%estimated ± 10.1 pp, low confidence
180GPT-4o0.0%estimated ± 10.1 pp, low confidence
181GPT-4o mini0.0%estimated ± 10.1 pp, low confidence
182Granite 4.2 3B0.0%estimated ± 10.1 pp, low confidence
183K-Exaone0.0%estimated ± 10.1 pp, low confidence
184Ling 2.6 Flash0.0%estimated ± 10.1 pp, low confidence
185Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 10.1 pp, low confidence
186North Mini Code0.0%estimated ± 10.1 pp, low confidence
187Solar Pro 30.0%estimated ± 10.1 pp, low confidence
188Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 10.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General