benchgap
Agentic · tools

WideResearch leaderboard

As of 2026-10-07, the highest measured score on WideResearch is 83.9% by Hy4 preview. 155 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5100.0%estimated ± 9.1 pp, low confidence
2Gemini 4 Argon100.0%estimated ± 9.1 pp, low confidence
3GPT-6.1 Sol100.0%estimated ± 9.1 pp, low confidence
4GPT-6 Sol100.0%estimated ± 9.1 pp, low confidence
5Grok 4.6100.0%estimated ± 9.1 pp, low confidence
6Grok 4.7100.0%estimated ± 9.1 pp, low confidence
7Qwen3.8 Max Preview100.0%estimated ± 9.1 pp, low confidence
8Muse Spark 1.298.8%estimated ± 9.1 pp, low confidence
9GPT-6 Luna97.1%estimated ± 9.1 pp, low confidence
10Gemini 3.8 Flash97.0%estimated ± 9.1 pp, low confidence
11Mistral Large 496.5%estimated ± 9.1 pp, low confidence
12Gemini 3.7 Flash95.1%estimated ± 9.1 pp, low confidence
13Grok 4.595.0%estimated ± 9.1 pp, low confidence
14Gemini 3.6 Flash90.3%estimated ± 9.1 pp, low confidence
15Claude Opus 5.587.9%estimated ± 3.4 pp, low confidence
16Claude Opus 586.5%estimated ± 3.4 pp, low confidence
17Claude Sonnet 5.586.4%estimated ± 3.4 pp, low confidence
18DeepSeek V4.1 Flash86.2%estimated ± 3.4 pp, low confidence
19Apodex 1.1 Mini86.0%estimated ± 9.1 pp, low confidence
20Claude Fable 5.185.6%estimated ± 3.9 pp, low confidence
21GLM-5.385.5%estimated ± 3.4 pp, low confidence
22Quasar 438B85.4%estimated ± 9.1 pp, low confidence
23Gemini 3.5 Flash84.7%estimated ± 5.6 pp, low confidence
24Muse Spark 1.384.5%estimated ± 1.2 pp, low confidence
25Ling 3.0 Flash VL84.4%estimated ± 9.1 pp, low confidence
26DeepSeek V4 Pro 081384.2%estimated ± 3.4 pp, low confidence
27Hy4 preview83.9%measured
28MiMo-V2.6-Pro83.9%estimated ± 1.2 pp, low confidence
29MiMo-V2.6-Flash83.7%estimated ± 1.2 pp, medium confidence
30Step 5 Preview83.2%estimated ± 1.2 pp, medium confidence
31Claude Haiku 5.582.8%estimated ± 3.4 pp, low confidence
32Claude Sonnet 582.8%estimated ± 3.4 pp, low confidence
33GPT-5.6 Sol82.8%estimated ± 4.6 pp, high confidence
34GPT-6 Astra82.7%estimated ± 3.4 pp, low confidence
35Ling 3.1 Flash82.7%estimated ± 1.5 pp, medium confidence
36Qwen3.8-Flash-Next82.5%estimated ± 1.2 pp, medium confidence
37Muse Spark 1.182.3%estimated ± 1.2 pp, medium confidence
38Apodex 1.182.1%estimated ± 3.4 pp, medium confidence
39GPT-5.5 Pro82.0%estimated ± 4.6 pp, high confidence
40Atria Dawn Preview81.9%measured
41Qwen3.8 Max81.9%measured
42Kimi K381.9%estimated ± 1.2 pp, medium confidence
43GPT-5.4 Pro81.7%estimated ± 4.6 pp, high confidence
44GLM-5.3-Flash81.6%estimated ± 3.4 pp, medium confidence
45Claude Mythos 581.2%estimated ± 4.6 pp, high confidence
46GPT-5.6 Terra81.0%estimated ± 4.6 pp, high confidence
47Kimi K2.680.8%measured
48Ornith-1.5-397B80.8%measured
49Qwen3.7 Max80.6%estimated ± 3.4 pp, medium confidence
50Claude Opus 4.7 (Adaptive)80.4%estimated ± 1.2 pp, medium confidence
51GPT-5.579.7%estimated ± 1.2 pp, medium confidence
52Claude Opus 4.879.7%estimated ± 4.6 pp, high confidence
53GLM-5.279.4%estimated ± 5.6 pp, low confidence
54Claude Opus 4.779.4%estimated ± 5.9 pp, low confidence
55MiniMax M379.3%estimated ± 4.6 pp, high confidence
56Hy379.2%estimated ± 9.1 pp, low confidence
57GPT-5.6 Luna79.2%estimated ± 4.6 pp, high confidence
58MiMo-V2.5-Pro78.9%estimated ± 5.9 pp, low confidence
59dots3-note Preview78.9%measured
60GPT-5.478.8%estimated ± 1.2 pp, medium confidence
61Claude Sonnet 4.678.4%estimated ± 1.2 pp, medium confidence
62Claude Opus 4.678.4%estimated ± 1.2 pp, medium confidence
63MiniMax M2.778.2%estimated ± 5.6 pp, low confidence
64GPT-5.277.8%estimated ± 1.2 pp, medium confidence
65Kimi K2.7 Code77.8%estimated ± 9.1 pp, low confidence
66GPT-5.3 Codex77.7%estimated ± 1.2 pp, medium confidence
67Qwen3.8-27B77.6%estimated ± 1.2 pp, medium confidence
68Gemini 3.1 Pro77.3%estimated ± 5.9 pp, low confidence
69Agents-A176.8%estimated ± 3.4 pp, medium confidence
70Beam76.6%estimated ± 4.6 pp, high confidence
71Inkling76.5%estimated ± 4.6 pp, high confidence
72Qwen3.6-27B76.5%estimated ± 5.9 pp, low confidence
73Step 3.7 Flash76.5%estimated ± 3.4 pp, medium confidence
74Claude Opus 4.576.4%measured
75Claude Sonnet 4.576.4%estimated ± 1.2 pp, medium confidence
76GPT-5.1-Codex76.1%estimated ± 1.2 pp, medium confidence
77GPT-5.4 mini76.0%estimated ± 5.6 pp, low confidence
78GPT-5.2-Codex76.0%estimated ± 1.2 pp, medium confidence
79Muse Spark75.7%estimated ± 9.1 pp, low confidence
80Claude 4.1 Opus75.1%estimated ± 1.2 pp, medium confidence
81DeepSeek V4 Flash 073175.0%estimated ± 3.4 pp, medium confidence
82Gemini 3.5 Flash-Lite74.8%estimated ± 9.1 pp, low confidence
83Qwen3.5 Plus74.4%estimated ± 1.2 pp, medium confidence
84Claude 4 Sonnet74.4%estimated ± 1.2 pp, medium confidence
85Qwen3.6 Plus74.3%measured
86Qwen3.5 397B74.0%measured
87Claude Haiku 4.573.8%estimated ± 1.2 pp, medium confidence
88Ling 3.0 Flash73.6%measured
89A.X K273.5%estimated ± 9.1 pp, low confidence
90Muse Glimmer 30B73.0%estimated ± 1.5 pp, low confidence
91Gemini 3 Flash72.8%estimated ± 1.2 pp, medium confidence
92Gemini 3 Pro72.8%estimated ± 1.2 pp, medium confidence
93Grok Build 0.172.8%estimated ± 5.9 pp, low confidence
94Kimi K2.572.7%measured
95Ling 3.0 Flash FP872.6%estimated ± 9.1 pp, low confidence
96Inkling-Small72.4%estimated ± 3.9 pp, medium confidence
97GPT-5 (high)72.2%estimated ± 1.2 pp, low confidence
98GLM-5.172.0%estimated ± 4.6 pp, high confidence
99Agents-A1-4B71.4%estimated ± 4.6 pp, high confidence
100Pokee-Isaac 28B71.4%estimated ± 12.4 pp, low confidence
101GPT-5.4 nano71.3%estimated ± 5.6 pp, low confidence
102Grok 4.1 Fast70.9%estimated ± 5.9 pp, low confidence
103MiMo-V2.570.4%estimated ± 5.9 pp, low confidence
104GLM-569.8%measured
105Qwen3.5-122B-A10B69.8%estimated ± 4.6 pp, high confidence
106Laguna S 2.169.1%estimated ± 3.9 pp, medium confidence
107Nemotron 3 Ultra68.8%estimated ± 3.4 pp, medium confidence
108Qwen3.5-27B68.2%estimated ± 4.6 pp, high confidence
109Qwen3.5-35B-A3B68.2%estimated ± 4.6 pp, high confidence
110Kimi K2.5 (Reasoning)67.9%estimated ± 4.6 pp, high confidence
111Ornith-1.5-35B-A3B67.8%measured
112Grok 4.366.2%estimated ± 5.9 pp, low confidence
113Qwen3 Max66.0%estimated ± 5.9 pp, low confidence
114Grok 463.4%estimated ± 5.9 pp, low confidence
115Gemini 2.5 Pro62.8%estimated ± 5.9 pp, low confidence
116GLM-4.762.5%estimated ± 4.6 pp, medium confidence
117Qwen3.7 Plus61.5%estimated ± 9.1 pp, low confidence
118GPT-5.161.2%estimated ± 5.9 pp, low confidence
119Solar Pro 460.6%estimated ± 4.6 pp, medium confidence
120LongCat-Flash-Lite-Sparse60.2%estimated ± 4.6 pp, medium confidence
121Qwen3.6-35B-A3B60.1%measured
122Ornith-1.5-9B59.5%measured
123MiniCPM5-2B57.4%estimated ± 9.1 pp, low confidence
124Mistral Medium 3.5 128B56.0%estimated ± 5.9 pp, low confidence
125Gemini 3.1 Flash-Lite54.2%estimated ± 5.9 pp, low confidence
126Grok 4.2054.0%estimated ± 5.9 pp, low confidence
127MiMo-V2-Flash51.9%estimated ± 9.1 pp, low confidence
128Nemotron 3.5 Lightning 30B A3B NVFP450.8%estimated ± 4.6 pp, medium confidence
129Hy3 Preview49.7%estimated ± 5.9 pp, low confidence
130MiMo-V2-Pro48.9%estimated ± 5.9 pp, low confidence
131Granite 4.2 30B47.8%estimated ± 9.1 pp, low confidence
132Gemma 4 26B A4B47.0%estimated ± 9.1 pp, low confidence
133Ling 3.0 Tiny46.7%estimated ± 9.1 pp, low confidence
134Gemma 4 31B44.2%estimated ± 5.9 pp, low confidence
135Command A+43.1%estimated ± 9.1 pp, low confidence
136Celeris-141.6%estimated ± 9.1 pp, low confidence
137DeepSeek V341.6%estimated ± 9.1 pp, low confidence
138DeepSeek V3 032441.6%estimated ± 9.1 pp, low confidence
139Gemma 3 27B41.6%estimated ± 9.1 pp, low confidence
140Gemma 4 12B41.6%estimated ± 9.1 pp, low confidence
141Gemma 4 E2B41.6%estimated ± 9.1 pp, low confidence
142Gemma 4 E4B41.6%estimated ± 9.1 pp, low confidence
143GPT-4.1 mini41.6%estimated ± 9.1 pp, low confidence
144GPT-4.1 nano41.6%estimated ± 9.1 pp, low confidence
145GPT-4o41.6%estimated ± 9.1 pp, low confidence
146GPT-4o mini41.6%estimated ± 9.1 pp, low confidence
147GPT-OSS 20B41.6%estimated ± 9.1 pp, low confidence
148Granite 4.2 3B41.6%estimated ± 9.1 pp, low confidence
149Granite 4.2 8B41.6%estimated ± 9.1 pp, low confidence
150K-Exaone41.6%estimated ± 9.1 pp, low confidence
151LFM2.5-2.6B41.6%estimated ± 9.1 pp, low confidence
152Ling 2.6 Flash41.6%estimated ± 9.1 pp, low confidence
153Llama 4 Maverick41.6%estimated ± 9.1 pp, low confidence
154Llama 4 Scout41.6%estimated ± 9.1 pp, low confidence
155Mercury 2.541.6%estimated ± 9.1 pp, low confidence
156Mistral Large 341.6%estimated ± 9.1 pp, low confidence
157Mistral Small 441.6%estimated ± 9.1 pp, low confidence
158Mistral Small 4 (Reasoning)41.6%estimated ± 9.1 pp, low confidence
159Nemotron 3 Nano 30B41.6%estimated ± 9.1 pp, low confidence
160Nemotron 3 Nano Omni 30B A3B41.6%estimated ± 9.1 pp, low confidence
161Nemotron 3 Super 100B41.6%estimated ± 9.1 pp, low confidence
162North Mini Code41.6%estimated ± 9.1 pp, low confidence
163Solar Pro 341.6%estimated ± 9.1 pp, low confidence
164Trinity-Large-Preview41.6%estimated ± 9.1 pp, low confidence
165Ultravox v0.6 Llama 3.3 70B41.6%estimated ± 9.1 pp, low confidence
166Trinity-Large-Thinking34.6%estimated ± 5.9 pp, low confidence
167GLM-5V-Turbo28.1%estimated ± 5.9 pp, low confidence
168GPT-OSS 120B24.0%estimated ± 5.9 pp, low confidence
169DeepSeek V3.223.9%estimated ± 5.9 pp, low confidence
170GPT-4.112.3%estimated ± 5.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General