benchgap
Agentic · tools

ResearchClawBench leaderboard

As of 2026-10-07, the highest measured score on ResearchClawBench is 21.1% by Claude Opus 4.8. 154 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Ornith-1.5-397B21.6%estimated ± 1.3 pp, medium confidence
2Claude Fable 5.121.5%estimated ± 2.0 pp, medium confidence
3Muse Spark 1.321.4%estimated ± 2.0 pp, medium confidence
4GLM-5.321.2%estimated ± 2.0 pp, medium confidence
5Grok 4.621.2%estimated ± 2.0 pp, medium confidence
6GPT-6 Astra21.1%estimated ± 2.0 pp, medium confidence
7Claude Opus 4.821.1%measured
8Claude Fable 521.1%estimated ± 2.0 pp, medium confidence
9GPT-5.6 Sol21.0%estimated ± 2.0 pp, medium confidence
10Qwen3.8 Max Preview21.0%estimated ± 2.0 pp, medium confidence
11K-EXAONE 2.020.8%estimated ± 1.3 pp, medium confidence
12Qwen3.8-27B20.7%estimated ± 2.0 pp, medium confidence
13Claude Opus 4.720.7%measured
14GLM-5.220.7%measured
15Ornith-1.0-397B20.6%estimated ± 1.3 pp, medium confidence
16Claude Opus 5.520.6%estimated ± 2.2 pp, low confidence
17Claude Sonnet 5.520.5%estimated ± 2.2 pp, low confidence
18Claude Sonnet 520.5%estimated ± 2.0 pp, medium confidence
19Muse Spark 1.220.5%estimated ± 2.0 pp, medium confidence
20GPT-5.6 Terra20.5%estimated ± 2.0 pp, medium confidence
21GPT-5.6 Luna20.3%estimated ± 2.0 pp, medium confidence
22Grok 4.520.3%estimated ± 2.0 pp, high confidence
23Grok 4.720.2%estimated ± 2.2 pp, low confidence
24Gemini 3.8 Flash20.2%estimated ± 2.0 pp, high confidence
25MiMo-V2.6-Pro20.1%estimated ± 2.2 pp, low confidence
26Atria Dawn Preview20.0%estimated ± 2.2 pp, low confidence
27Qwen3.8-Flash-Next19.9%estimated ± 2.2 pp, low confidence
28Gemini 4 Argon19.9%estimated ± 2.2 pp, low confidence
29Ling 3.1 Flash19.9%estimated ± 2.2 pp, low confidence
30Claude Opus 4.619.9%measured
31Muse Spark 1.119.9%estimated ± 1.8 pp, low confidence
32MiMo-V2.6-Flash19.9%estimated ± 2.2 pp, low confidence
33DeepSeek V4.1 Flash19.8%estimated ± 2.2 pp, low confidence
34MiniMax M319.8%measured
35GPT-6.1 Sol19.8%estimated ± 2.2 pp, low confidence
36dots3-note Preview19.7%estimated ± 1.3 pp, high confidence
37Claude Opus 519.7%estimated ± 1.8 pp, low confidence
38GPT-5.5 Pro19.6%estimated ± 2.2 pp, low confidence
39Step 5 Preview19.6%estimated ± 1.8 pp, low confidence
40Gemini 3.7 Flash19.6%estimated ± 2.0 pp, high confidence
41GPT-5.4 Pro19.5%estimated ± 2.2 pp, low confidence
42GPT-6 Sol19.5%estimated ± 2.2 pp, low confidence
43Kimi K319.5%estimated ± 1.8 pp, low confidence
44Ornith-1.5-35B-A3B19.5%estimated ± 1.3 pp, high confidence
45Qwen3.6-27B19.5%estimated ± 1.3 pp, high confidence
46Hy4 preview19.5%estimated ± 1.8 pp, low confidence
47Claude Mythos 519.3%estimated ± 2.2 pp, low confidence
48GPT-6 Luna19.2%estimated ± 2.2 pp, medium confidence
49Mistral Large 419.2%estimated ± 2.2 pp, medium confidence
50Quasar 438B19.1%estimated ± 2.0 pp, high confidence
51Inkling-Small19.1%estimated ± 1.8 pp, medium confidence
52Beam19.0%estimated ± 1.8 pp, medium confidence
53Claude Opus 4.7 (Adaptive)18.8%estimated ± 1.8 pp, medium confidence
54Ornith-1.0-35B18.8%estimated ± 1.3 pp, high confidence
55Kimi K2.7 Code18.7%estimated ± 1.8 pp, medium confidence
56Gemini 3.6 Flash18.7%estimated ± 2.0 pp, high confidence
57Qwen3.7 Max18.7%measured
58Muse Glimmer 30B18.7%estimated ± 1.8 pp, medium confidence
59Qwen3.6-35B-A3B18.6%estimated ± 1.3 pp, high confidence
60Inkling18.5%estimated ± 1.8 pp, medium confidence
61DeepSeek V4 Pro 081318.5%estimated ± 1.8 pp, medium confidence
62Claude Sonnet 4.618.3%estimated ± 1.3 pp, high confidence
63GLM-5.118.2%measured
64Step 3.7 Flash18.2%estimated ± 1.3 pp, high confidence
65DeepSeek V4 Flash 073118.0%estimated ± 1.8 pp, medium confidence
66Ornith-1.5-9B18.0%estimated ± 1.3 pp, high confidence
67Gemini 3.5 Flash18.0%measured
68Kimi K2.618.0%measured
69Qwen3.6 Plus18.0%measured
70Gemini 3 Pro18.0%estimated ± 2.2 pp, medium confidence
71Apodex 1.117.9%estimated ± 2.2 pp, medium confidence
72Apodex 1.1 Mini17.9%estimated ± 2.2 pp, medium confidence
73Hy317.9%estimated ± 2.0 pp, high confidence
74Hy3 Preview17.9%estimated ± 2.0 pp, high confidence
75Ling 3.0 Flash17.7%estimated ± 1.8 pp, medium confidence
76Ling 3.0 Flash VL17.7%estimated ± 2.2 pp, medium confidence
77LLaDA2.2-flash17.5%estimated ± 1.3 pp, high confidence
78MiMo-V2.5-Pro17.4%estimated ± 1.3 pp, high confidence
79Muse Spark17.4%estimated ± 1.3 pp, high confidence
80Agents-A117.3%estimated ± 2.2 pp, medium confidence
81Solar Pro 417.3%estimated ± 1.8 pp, medium confidence
82Ornith-1.0-9B17.2%estimated ± 1.3 pp, high confidence
83LFM2.5-2.6B17.1%estimated ± 1.3 pp, high confidence
84Qwen3.7 Plus17.1%estimated ± 1.3 pp, high confidence
85GPT-5.3 Codex17.0%estimated ± 2.2 pp, medium confidence
86Solar Open 217.0%estimated ± 1.8 pp, medium confidence
87GPT-5.517.0%measured
88GPT-5.4 mini17.0%estimated ± 1.8 pp, medium confidence
89Nemotron 3 Ultra16.9%estimated ± 2.0 pp, high confidence
90MiMo-V2.516.9%measured
91GPT-5.4 nano16.8%estimated ± 1.8 pp, medium confidence
92Ling 3.0 Flash FP816.7%estimated ± 2.0 pp, high confidence
93Claude Opus 4.6 (Adaptive)16.5%estimated ± 2.6 pp, medium confidence
94Claude Opus 4.516.3%estimated ± 1.3 pp, high confidence
95GPT-5.2-Codex16.1%estimated ± 2.2 pp, medium confidence
96A.X K215.8%estimated ± 2.2 pp, medium confidence
97GLM-515.8%estimated ± 1.3 pp, high confidence
98Agents-A1-4B15.8%estimated ± 2.2 pp, medium confidence
99GPT-5.1-Codex15.8%estimated ± 2.2 pp, medium confidence
100LongCat-Flash-Lite-Sparse15.8%estimated ± 1.8 pp, medium confidence
101Grok Build 0.115.7%estimated ± 2.2 pp, medium confidence
102LLaDA2.2-mini15.7%estimated ± 1.3 pp, high confidence
103GPT-5.215.6%estimated ± 2.2 pp, medium confidence
104Claude Sonnet 4.515.6%estimated ± 2.2 pp, medium confidence
105Grok 4.1 Fast15.4%estimated ± 2.2 pp, medium confidence
106GLM-5-Turbo15.4%estimated ± 1.3 pp, high confidence
107GPT-5.415.3%measured
108MiMo-V2-Pro15.3%measured
109GPT-5 (high)15.3%estimated ± 2.2 pp, medium confidence
110Gemini 3.5 Flash-Lite15.0%estimated ± 2.0 pp, high confidence
111Qwen3 Max14.9%estimated ± 2.2 pp, medium confidence
112GLM-5V-Turbo14.8%estimated ± 1.3 pp, high confidence
113Qwen3.5-27B14.8%estimated ± 2.2 pp, medium confidence
114Qwen3.5-35B-A3B14.8%estimated ± 2.2 pp, medium confidence
115Kimi K2.5 (Reasoning)14.7%estimated ± 2.2 pp, medium confidence
116Grok 414.7%estimated ± 2.2 pp, medium confidence
117Claude 4 Sonnet14.2%estimated ± 2.2 pp, medium confidence
118Qwen3.5 397B14.2%measured
119Gemini 3.1 Flash-Lite14.0%estimated ± 2.2 pp, medium confidence
120Grok 4.2014.0%estimated ± 2.2 pp, medium confidence
121Kimi K2.514.0%measured
122GPT-5.113.9%estimated ± 2.2 pp, medium confidence
123Gemini 3 Flash13.7%estimated ± 1.3 pp, medium confidence
124MiniMax M2.713.5%estimated ± 1.3 pp, medium confidence
125Grok 4.113.5%measured
126Gemini 3.1 Pro13.3%measured
127GLM-4.713.0%estimated ± 2.2 pp, low confidence
128MiMo-V2-Omni12.6%estimated ± 1.3 pp, medium confidence
129Grok 4.312.4%measured
130GPT-4.112.0%estimated ± 2.2 pp, low confidence
131Qwen3.5-122B-A10B11.8%estimated ± 2.0 pp, medium confidence
132Mistral Medium 3.5 128B11.6%estimated ± 2.0 pp, medium confidence
133DeepSeek V3.211.3%estimated ± 1.3 pp, medium confidence
134MiniCPM5-2B11.2%estimated ± 2.2 pp, low confidence
135Gemma 4 31B9.6%estimated ± 2.0 pp, medium confidence
136GPT-OSS 120B9.1%estimated ± 2.0 pp, medium confidence
137Nemotron 3.5 Lightning 30B A3B NVFP49.0%estimated ± 2.0 pp, medium confidence
138MiMo-V2-Flash8.5%estimated ± 2.2 pp, low confidence
139Granite 4.2 8B6.3%estimated ± 2.0 pp, medium confidence
140Command A+6.2%estimated ± 2.0 pp, medium confidence
141Gemini 2.5 Pro6.1%estimated ± 2.0 pp, medium confidence
142Granite 4.2 30B5.8%estimated ± 2.2 pp, low confidence
143Gemma 4 26B A4B5.2%estimated ± 2.2 pp, low confidence
144Ling 3.0 Tiny5.0%estimated ± 2.2 pp, low confidence
145Mistral Large 34.6%estimated ± 2.0 pp, medium confidence
146Mistral Small 42.9%estimated ± 2.0 pp, medium confidence
147Mistral Small 4 (Reasoning)2.9%estimated ± 2.0 pp, medium confidence
148GPT-OSS 20B2.8%estimated ± 2.0 pp, medium confidence
149Trinity-Large-Preview2.4%estimated ± 2.0 pp, medium confidence
150Trinity-Large-Thinking2.4%estimated ± 2.0 pp, medium confidence
151Nemotron 3 Nano 30B2.1%estimated ± 2.0 pp, medium confidence
152DeepSeek V31.7%estimated ± 2.0 pp, medium confidence
153Nemotron 3 Super 100B1.6%estimated ± 1.3 pp, medium confidence
154Celeris-11.4%estimated ± 2.0 pp, medium confidence
155Llama 4 Maverick1.3%estimated ± 2.0 pp, medium confidence
156Llama 4 Scout1.2%estimated ± 2.0 pp, medium confidence
157Gemma 3 27B0.3%estimated ± 2.0 pp, medium confidence
158DeepSeek V3 03240.0%estimated ± 2.2 pp, low confidence
159Gemma 4 12B0.0%estimated ± 2.2 pp, low confidence
160Gemma 4 E2B0.0%estimated ± 2.2 pp, low confidence
161Gemma 4 E4B0.0%estimated ± 2.2 pp, low confidence
162GPT-4.1 mini0.0%estimated ± 2.2 pp, low confidence
163GPT-4.1 nano0.0%estimated ± 2.2 pp, low confidence
164GPT-4o0.0%estimated ± 2.2 pp, low confidence
165GPT-4o mini0.0%estimated ± 2.2 pp, low confidence
166Granite 4.2 3B0.0%estimated ± 2.2 pp, low confidence
167K-Exaone0.0%estimated ± 2.2 pp, low confidence
168Ling 2.6 Flash0.0%estimated ± 2.2 pp, low confidence
169Mercury 2.50.0%estimated ± 2.2 pp, low confidence
170Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 2.2 pp, low confidence
171North Mini Code0.0%estimated ± 2.2 pp, low confidence
172Solar Pro 30.0%estimated ± 2.2 pp, low confidence
173Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 2.2 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General