benchgap
Agentic · tools

QwenClawBench leaderboard

As of 2026-10-07, the highest measured score on QwenClawBench is 64.3% by Qwen3.7 Max. 128 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5100.0%estimated ± 6.0 pp, low confidence
2Claude Fable 5.1100.0%estimated ± 6.0 pp, low confidence
3Claude Opus 5.5100.0%estimated ± 6.0 pp, low confidence
4Claude Sonnet 5100.0%estimated ± 6.0 pp, low confidence
5Claude Sonnet 5.5100.0%estimated ± 6.0 pp, low confidence
6DeepSeek V4.1 Flash100.0%estimated ± 6.0 pp, low confidence
7Gemini 3.6 Flash100.0%estimated ± 6.0 pp, low confidence
8Gemini 3.7 Flash100.0%estimated ± 6.0 pp, low confidence
9Gemini 3.8 Flash100.0%estimated ± 6.0 pp, low confidence
10Gemini 4 Argon100.0%estimated ± 6.0 pp, low confidence
11GLM-5.3100.0%estimated ± 6.0 pp, low confidence
12GPT-5.6 Luna100.0%estimated ± 6.0 pp, low confidence
13GPT-5.6 Sol100.0%estimated ± 6.0 pp, low confidence
14GPT-5.6 Terra100.0%estimated ± 6.0 pp, low confidence
15GPT-6.1 Sol100.0%estimated ± 6.0 pp, low confidence
16GPT-6 Astra100.0%estimated ± 6.0 pp, low confidence
17GPT-6 Luna100.0%estimated ± 6.0 pp, low confidence
18GPT-6 Sol100.0%estimated ± 6.0 pp, low confidence
19Grok 4.5100.0%estimated ± 6.0 pp, low confidence
20Grok 4.6100.0%estimated ± 6.0 pp, low confidence
21Grok 4.7100.0%estimated ± 6.0 pp, low confidence
22Hy3 Preview100.0%estimated ± 6.0 pp, low confidence
23Ling 3.1 Flash100.0%estimated ± 6.0 pp, low confidence
24MiMo-V2.6-Flash100.0%estimated ± 6.0 pp, low confidence
25MiMo-V2.6-Pro100.0%estimated ± 6.0 pp, low confidence
26Mistral Large 4100.0%estimated ± 6.0 pp, low confidence
27Muse Spark 1.2100.0%estimated ± 6.0 pp, low confidence
28Muse Spark 1.3100.0%estimated ± 6.0 pp, low confidence
29Qwen3.8-27B100.0%estimated ± 6.0 pp, low confidence
30Qwen3.8-Flash-Next100.0%estimated ± 6.0 pp, low confidence
31Qwen3.8 Max Preview100.0%estimated ± 6.0 pp, low confidence
32Apodex 1.199.4%estimated ± 6.0 pp, low confidence
33Apodex 1.1 Mini99.4%estimated ± 6.0 pp, low confidence
34Quasar 438B97.9%estimated ± 6.0 pp, low confidence
35Ling 3.0 Flash VL89.8%estimated ± 6.0 pp, low confidence
36Nemotron 3 Ultra88.5%estimated ± 6.0 pp, low confidence
37Muse Spark 1.173.0%estimated ± 2.6 pp, medium confidence
38Claude Opus 570.9%estimated ± 2.6 pp, medium confidence
39Step 5 Preview70.7%estimated ± 2.6 pp, medium confidence
40Kimi K369.4%estimated ± 2.6 pp, medium confidence
41Hy4 preview69.0%estimated ± 2.6 pp, medium confidence
42Gemini 3.5 Flash68.9%estimated ± 2.6 pp, medium confidence
43Claude Opus 4.867.7%estimated ± 2.6 pp, medium confidence
44Ornith-1.5-397B65.9%estimated ± 2.6 pp, medium confidence
45Inkling-Small65.6%estimated ± 2.6 pp, medium confidence
46Beam65.0%estimated ± 2.6 pp, medium confidence
47Qwen3.7 Max64.3%measured
48Claude Opus 4.7 (Adaptive)63.9%estimated ± 2.6 pp, medium confidence
49GLM-5.263.6%estimated ± 2.6 pp, medium confidence
50Kimi K2.7 Code63.0%estimated ± 2.6 pp, high confidence
51Muse Glimmer 30B62.7%estimated ± 2.6 pp, high confidence
52GPT-5.562.6%estimated ± 2.6 pp, high confidence
53MiniMax M361.9%estimated ± 2.6 pp, high confidence
54Inkling61.8%estimated ± 2.6 pp, high confidence
55Qwen3.7 Plus61.8%measured
56DeepSeek V4 Pro 081361.5%estimated ± 2.6 pp, high confidence
57MiMo-V2.5-Pro60.7%estimated ± 6.0 pp, low confidence
58GLM-5.160.5%estimated ± 2.6 pp, high confidence
59GPT-5.459.8%estimated ± 2.6 pp, high confidence
60Ornith-1.5-35B-A3B59.6%estimated ± 2.6 pp, high confidence
61DeepSeek V4 Flash 073159.1%estimated ± 2.6 pp, high confidence
62Qwen 3.6 Max (preview)59.0%measured
63Ling 3.0 Flash57.6%estimated ± 2.6 pp, high confidence
64Qwen3.6 Plus57.2%measured
65Solar Pro 456.2%estimated ± 2.6 pp, high confidence
66Grok 4.356.1%estimated ± 6.0 pp, low confidence
67Hy355.9%estimated ± 6.0 pp, low confidence
68Step 3.7 Flash55.9%estimated ± 6.0 pp, low confidence
69MiniMax M2.755.9%estimated ± 6.0 pp, low confidence
70Muse Spark55.9%estimated ± 6.0 pp, low confidence
71A.X K255.9%estimated ± 6.0 pp, low confidence
72Celeris-155.9%estimated ± 6.0 pp, low confidence
73Command A+55.9%estimated ± 6.0 pp, low confidence
74DeepSeek V355.9%estimated ± 6.0 pp, low confidence
75DeepSeek V3 032455.9%estimated ± 6.0 pp, low confidence
76Gemini 2.5 Pro55.9%estimated ± 6.0 pp, low confidence
77Gemini 3.1 Pro55.9%estimated ± 6.0 pp, low confidence
78Gemini 3.5 Flash-Lite55.9%estimated ± 6.0 pp, low confidence
79Gemma 3 27B55.9%estimated ± 6.0 pp, low confidence
80Gemma 4 12B55.9%estimated ± 6.0 pp, low confidence
81Gemma 4 26B A4B55.9%estimated ± 6.0 pp, low confidence
82Gemma 4 31B55.9%estimated ± 6.0 pp, low confidence
83Gemma 4 E2B55.9%estimated ± 6.0 pp, low confidence
84Gemma 4 E4B55.9%estimated ± 6.0 pp, low confidence
85GPT-4.1 mini55.9%estimated ± 6.0 pp, low confidence
86GPT-4.1 nano55.9%estimated ± 6.0 pp, low confidence
87GPT-4o55.9%estimated ± 6.0 pp, low confidence
88GPT-4o mini55.9%estimated ± 6.0 pp, low confidence
89GPT-5.155.9%estimated ± 6.0 pp, low confidence
90GPT-5 (high)55.9%estimated ± 6.0 pp, low confidence
91GPT-OSS 120B55.9%estimated ± 6.0 pp, low confidence
92GPT-OSS 20B55.9%estimated ± 6.0 pp, low confidence
93Granite 4.2 30B55.9%estimated ± 6.0 pp, low confidence
94Granite 4.2 3B55.9%estimated ± 6.0 pp, low confidence
95Granite 4.2 8B55.9%estimated ± 6.0 pp, low confidence
96K-Exaone55.9%estimated ± 6.0 pp, low confidence
97Kimi K2.5 (Reasoning)55.9%estimated ± 6.0 pp, low confidence
98LFM2.5-2.6B55.9%estimated ± 6.0 pp, low confidence
99Ling 2.6 Flash55.9%estimated ± 6.0 pp, low confidence
100Ling 3.0 Flash FP855.9%estimated ± 6.0 pp, low confidence
101Ling 3.0 Tiny55.9%estimated ± 6.0 pp, low confidence
102Llama 4 Maverick55.9%estimated ± 6.0 pp, low confidence
103Llama 4 Scout55.9%estimated ± 6.0 pp, low confidence
104Mercury 2.555.9%estimated ± 6.0 pp, low confidence
105MiMo-V2-Flash55.9%estimated ± 6.0 pp, low confidence
106MiniCPM5-2B55.9%estimated ± 6.0 pp, low confidence
107Mistral Large 355.9%estimated ± 6.0 pp, low confidence
108Mistral Medium 3.5 128B55.9%estimated ± 6.0 pp, low confidence
109Mistral Small 455.9%estimated ± 6.0 pp, low confidence
110Mistral Small 4 (Reasoning)55.9%estimated ± 6.0 pp, low confidence
111Nemotron 3.5 Lightning 30B A3B NVFP455.9%estimated ± 6.0 pp, low confidence
112Nemotron 3 Nano 30B55.9%estimated ± 6.0 pp, low confidence
113Nemotron 3 Nano Omni 30B A3B55.9%estimated ± 6.0 pp, low confidence
114Nemotron 3 Super 100B55.9%estimated ± 6.0 pp, low confidence
115North Mini Code55.9%estimated ± 6.0 pp, low confidence
116Qwen3.5-122B-A10B55.9%estimated ± 6.0 pp, low confidence
117Solar Pro 355.9%estimated ± 6.0 pp, low confidence
118Trinity-Large-Preview55.9%estimated ± 6.0 pp, low confidence
119Trinity-Large-Thinking55.9%estimated ± 6.0 pp, low confidence
120Ultravox v0.6 Llama 3.3 70B55.9%estimated ± 6.0 pp, low confidence
121Solar Open 255.4%estimated ± 2.6 pp, high confidence
122GPT-5.4 mini55.2%estimated ± 2.6 pp, high confidence
123GPT-5.4 nano54.9%estimated ± 2.6 pp, high confidence
124Kimi K2.654.9%estimated ± 2.6 pp, high confidence
125Ornith-1.5-9B54.6%estimated ± 2.6 pp, high confidence
126Kimi K2.554.3%measured
127GLM-554.1%measured
128LLaDA2.2-flash53.7%estimated ± 2.6 pp, high confidence
129Qwen3.6-27B53.4%measured
130Qwen3.6-35B-A3B52.6%measured
131Claude Opus 4.552.3%measured
132Agents-A1-4B52.0%estimated ± 2.0 pp, medium confidence
133Agents-A152.0%estimated ± 2.0 pp, medium confidence
134Claude Sonnet 4.552.0%estimated ± 2.0 pp, low confidence
135DeepSeek V3.252.0%estimated ± 2.0 pp, low confidence
136GLM-4.752.0%estimated ± 2.0 pp, low confidence
137LongCat-Flash-Lite-Sparse52.0%estimated ± 2.0 pp, low confidence
138Qwen3.5 397B51.8%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General