benchgap
Agentic · tools

DeepSearchQA leaderboard

As of 2026-10-07, the highest measured score on DeepSearchQA is 96.0% by Atria Dawn Preview. 148 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Fugu Cyber100.0%estimated ± 11.5 pp, low confidence
2Atria Dawn Preview96.0%measured
3Claude Opus 595.0%measured
4Kimi K395.0%measured
5Claude Fable 5.194.3%estimated ± 4.5 pp, medium confidence
6GLM-5.393.8%estimated ± 4.5 pp, high confidence
7Grok 4.693.8%estimated ± 4.5 pp, high confidence
8GPT-6 Astra93.6%estimated ± 4.5 pp, high confidence
9Claude Fable 593.5%estimated ± 4.5 pp, high confidence
10GPT-5.6 Sol93.4%estimated ± 4.5 pp, high confidence
11Qwen3.8 Max Preview93.3%estimated ± 4.5 pp, high confidence
12DeepSeek V4 Pro 081393.3%estimated ± 4.5 pp, high confidence
13Claude Opus 4.893.1%measured
14Qwen3.8-27B92.9%estimated ± 4.5 pp, high confidence
15Step 3.7 Flash92.8%measured
16Claude Sonnet 592.5%estimated ± 4.5 pp, high confidence
17Kimi K2.692.5%measured
18Muse Spark 1.292.5%estimated ± 4.5 pp, high confidence
19GPT-5.6 Terra92.4%estimated ± 4.5 pp, high confidence
20Apodex 1.192.4%measured
21GPT-5.6 Luna92.2%estimated ± 4.5 pp, high confidence
22Grok 4.592.1%estimated ± 4.5 pp, high confidence
23GPT-5.5 Pro92.1%estimated ± 8.1 pp, low confidence
24dots3-note Preview92.1%measured
25DeepSeek V4 Flash 073192.1%estimated ± 4.5 pp, high confidence
26Gemini 3.8 Flash91.9%estimated ± 4.5 pp, high confidence
27Claude Opus 4.7 (Adaptive)91.6%estimated ± 4.5 pp, high confidence
28GLM-5.291.6%estimated ± 4.5 pp, high confidence
29GPT-5.4 Pro91.6%estimated ± 8.1 pp, low confidence
30GPT-5.591.1%estimated ± 4.5 pp, high confidence
31Step 5 Preview91.1%estimated ± 8.1 pp, low confidence
32Gemini 3.7 Flash90.9%estimated ± 4.5 pp, high confidence
33Claude Mythos 590.6%estimated ± 8.1 pp, low confidence
34Claude Opus 5.590.5%estimated ± 8.4 pp, low confidence
35Claude Sonnet 5.590.5%estimated ± 8.4 pp, low confidence
36Grok 4.790.5%estimated ± 8.4 pp, medium confidence
37MiMo-V2.6-Pro90.5%estimated ± 8.4 pp, medium confidence
38Qwen3.8-Flash-Next90.4%estimated ± 8.4 pp, medium confidence
39Gemini 4 Argon90.4%estimated ± 8.4 pp, medium confidence
40Ling 3.1 Flash90.4%estimated ± 8.4 pp, medium confidence
41MiMo-V2.6-Flash90.4%estimated ± 8.4 pp, medium confidence
42DeepSeek V4.1 Flash90.4%estimated ± 8.4 pp, medium confidence
43GPT-6.1 Sol90.4%estimated ± 8.4 pp, medium confidence
44GPT-6 Sol90.3%estimated ± 8.4 pp, medium confidence
45Hy4 preview90.2%estimated ± 10.9 pp, low confidence
46GPT-6 Luna90.2%estimated ± 8.4 pp, medium confidence
47Mistral Large 490.1%estimated ± 8.4 pp, medium confidence
48Gemini 3.8 Flash Cyber90.0%estimated ± 11.5 pp, low confidence
49Quasar 438B89.9%estimated ± 4.5 pp, high confidence
50Ornith-1.5-397B89.6%estimated ± 8.1 pp, low confidence
51Muse Spark 1.389.4%measured
52MiniMax M389.4%estimated ± 4.5 pp, high confidence
53Gemini 3.6 Flash89.2%estimated ± 4.5 pp, high confidence
54Qwen3.8 Max88.9%estimated ± 8.5 pp, low confidence
55Apodex 1.1 Mini88.8%estimated ± 8.4 pp, medium confidence
56Ling 3.0 Flash VL88.4%estimated ± 8.4 pp, medium confidence
57Gemini 3.5 Flash88.1%estimated ± 4.5 pp, high confidence
58Hy387.4%estimated ± 4.5 pp, high confidence
59Hy3 Preview87.4%estimated ± 4.5 pp, high confidence
60GLM-5.187.2%estimated ± 4.5 pp, high confidence
61Inkling-Small87.1%estimated ± 4.5 pp, high confidence
62Inkling86.8%estimated ± 4.5 pp, high confidence
63Holo3-35B-A3B86.8%estimated ± 8.5 pp, low confidence
64Qwen3.7 Max86.6%estimated ± 4.5 pp, high confidence
65MiMo-V2.5-Pro85.9%estimated ± 4.5 pp, high confidence
66Kimi K2.7 Code85.8%estimated ± 4.5 pp, high confidence
67Nemotron 3 Ultra85.4%estimated ± 4.5 pp, high confidence
68Ling 3.0 Flash84.9%estimated ± 4.5 pp, high confidence
69Ling 3.0 Flash FP884.9%estimated ± 4.5 pp, high confidence
70Muse Spark 1.184.9%measured
71Holo3-122B-A10B84.5%estimated ± 8.5 pp, low confidence
72Qwen3.6 Plus84.5%estimated ± 8.4 pp, medium confidence
73Qwen3.6-27B84.3%estimated ± 4.5 pp, high confidence
74Qwen3.7 Plus84.1%estimated ± 4.5 pp, high confidence
75GPT-5.4 mini84.0%estimated ± 4.5 pp, high confidence
76UI-Mate-27B83.4%estimated ± 8.5 pp, low confidence
77A.X K283.3%estimated ± 8.4 pp, medium confidence
78Agents-A182.7%estimated ± 8.1 pp, low confidence
79GPT-5.4 nano82.6%estimated ± 4.5 pp, high confidence
80Grok 4.382.2%estimated ± 4.5 pp, high confidence
81MiniMax M2.781.8%estimated ± 4.5 pp, high confidence
82GPT-5 (high)81.2%estimated ± 8.4 pp, medium confidence
83Gemini 3.5 Flash-Lite80.9%estimated ± 4.5 pp, high confidence
84MiMo-V2.580.5%estimated ± 10.2 pp, low confidence
85Beam80.1%measured
86Qwen3.6-35B-A3B80.0%estimated ± 4.5 pp, high confidence
87Ornith-1.5-35B-A3B79.2%estimated ± 8.1 pp, low confidence
88Agents-A1-4B78.9%estimated ± 8.1 pp, low confidence
89GPT-5.278.6%estimated ± 8.1 pp, low confidence
90GPT-5.1-Codex78.3%estimated ± 10.9 pp, low confidence
91GPT-5.2-Codex78.3%estimated ± 10.9 pp, low confidence
92Qwen3.5 397B77.5%estimated ± 8.1 pp, low confidence
93Qwen3.5-27B77.2%estimated ± 8.1 pp, low confidence
94Qwen3.5-35B-A3B77.2%estimated ± 8.1 pp, low confidence
95Claude Opus 4.777.2%estimated ± 7.5 pp, low confidence
96Kimi K2.5 (Reasoning)77.1%estimated ± 8.1 pp, low confidence
97Kimi K2.577.1%measured
98MiMo-V2-Pro77.0%estimated ± 10.2 pp, low confidence
99Claude 4.1 Opus76.9%estimated ± 10.9 pp, low confidence
100Gemini 3.5 Flash Cyber76.8%estimated ± 11.5 pp, low confidence
101Claude Mythos Preview76.8%estimated ± 11.5 pp, low confidence
102GLM-576.8%estimated ± 11.5 pp, low confidence
103Ornith-1.5-9B76.3%estimated ± 8.1 pp, low confidence
104Claude Opus 4.576.2%estimated ± 8.5 pp, low confidence
105UI-Mate-9B76.1%estimated ± 8.5 pp, low confidence
106Qwen3.5 Plus75.7%estimated ± 10.9 pp, low confidence
107Claude 4 Sonnet75.7%estimated ± 10.9 pp, low confidence
108GLM-4.775.7%estimated ± 8.1 pp, low confidence
109Solar Pro 475.4%estimated ± 8.1 pp, low confidence
110LongCat-Flash-Lite-Sparse75.3%estimated ± 8.1 pp, low confidence
111GPT-5.175.3%estimated ± 8.4 pp, medium confidence
112GPT-5.3 Codex75.1%estimated ± 8.5 pp, low confidence
113Claude Haiku 4.574.9%estimated ± 10.9 pp, low confidence
114Muse Spark74.8%measured
115Muse Glimmer 30B74.6%measured
116Claude Opus 4.673.7%measured
117GPT-5.473.6%measured
118Claude Sonnet 4.673.4%estimated ± 7.5 pp, low confidence
119Gemini 3 Flash73.4%estimated ± 10.9 pp, low confidence
120Gemini 3 Pro73.4%estimated ± 10.9 pp, low confidence
121Claude Sonnet 4.572.6%estimated ± 8.5 pp, low confidence
122Grok 4.172.6%estimated ± 10.2 pp, low confidence
123Qwen3.5-122B-A10B71.8%estimated ± 4.5 pp, medium confidence
124Mistral Medium 3.5 128B71.1%estimated ± 4.5 pp, medium confidence
125Gemini 3.1 Pro69.7%measured
126Gemma 4 31B63.9%estimated ± 4.5 pp, medium confidence
127MiniCPM5-2B63.2%estimated ± 8.4 pp, medium confidence
128Grok 4.2062.8%measured
129GPT-OSS 120B62.0%estimated ± 4.5 pp, medium confidence
130Nemotron 3.5 Lightning 30B A3B NVFP461.7%estimated ± 4.5 pp, medium confidence
131MiMo-V2-Flash53.0%estimated ± 8.4 pp, medium confidence
132Nemotron 3 Super 100B51.9%estimated ± 4.5 pp, medium confidence
133Granite 4.2 8B49.2%estimated ± 4.5 pp, medium confidence
134Command A+48.7%estimated ± 4.5 pp, medium confidence
135Gemini 2.5 Pro48.1%estimated ± 4.5 pp, medium confidence
136Granite 4.2 30B45.4%estimated ± 8.4 pp, medium confidence
137Gemma 4 26B A4B44.0%estimated ± 8.4 pp, medium confidence
138Ling 3.0 Tiny43.4%estimated ± 8.4 pp, medium confidence
139Mistral Large 338.9%estimated ± 4.5 pp, medium confidence
140DeepSeek V3 032434.5%estimated ± 8.4 pp, medium confidence
141Gemma 4 12B34.5%estimated ± 8.4 pp, medium confidence
142Gemma 4 E2B34.5%estimated ± 8.4 pp, medium confidence
143Gemma 4 E4B34.5%estimated ± 8.4 pp, medium confidence
144GPT-4.1 mini34.5%estimated ± 8.4 pp, medium confidence
145GPT-4.1 nano34.5%estimated ± 8.4 pp, medium confidence
146GPT-4o34.5%estimated ± 8.4 pp, medium confidence
147GPT-4o mini34.5%estimated ± 8.4 pp, medium confidence
148Granite 4.2 3B34.5%estimated ± 8.4 pp, medium confidence
149K-Exaone34.5%estimated ± 8.4 pp, medium confidence
150LFM2.5-2.6B34.5%estimated ± 8.4 pp, medium confidence
151Ling 2.6 Flash34.5%estimated ± 8.4 pp, medium confidence
152Nemotron 3 Nano Omni 30B A3B34.5%estimated ± 8.4 pp, medium confidence
153North Mini Code34.5%estimated ± 8.4 pp, medium confidence
154Solar Pro 334.5%estimated ± 8.4 pp, medium confidence
155Ultravox v0.6 Llama 3.3 70B34.5%estimated ± 8.4 pp, medium confidence
156Mercury 2.534.0%measured
157Mistral Small 427.3%estimated ± 4.5 pp, medium confidence
158Mistral Small 4 (Reasoning)27.3%estimated ± 4.5 pp, medium confidence
159GPT-OSS 20B26.7%estimated ± 4.5 pp, medium confidence
160Trinity-Large-Preview23.2%estimated ± 4.5 pp, medium confidence
161Trinity-Large-Thinking23.2%estimated ± 4.5 pp, medium confidence
162Nemotron 3 Nano 30B20.7%estimated ± 4.5 pp, medium confidence
163DeepSeek V317.0%estimated ± 4.5 pp, medium confidence
164Celeris-114.5%estimated ± 4.5 pp, medium confidence
165Llama 4 Maverick13.9%estimated ± 4.5 pp, medium confidence
166Llama 4 Scout12.5%estimated ± 4.5 pp, medium confidence
167Gemma 3 27B3.5%estimated ± 4.5 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General