benchgap
Agentic · tools

Claw-Eval leaderboard

As of 2026-10-07, the highest measured score on Claw-Eval is 81.4% by Ornith-1.5-397B. 116 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.186.8%estimated ± 13.6 pp, low confidence
2GLM-5.385.3%estimated ± 13.6 pp, low confidence
3Grok 4.685.3%estimated ± 13.6 pp, low confidence
4Claude Fable 584.5%estimated ± 13.6 pp, low confidence
5Qwen3.8 Max Preview84.0%estimated ± 13.6 pp, low confidence
6GPT-5.6 Sol81.7%estimated ± 7.9 pp, low confidence
7GPT-6 Astra81.7%estimated ± 7.9 pp, low confidence
8GPT-5.5 Pro81.7%estimated ± 7.9 pp, low confidence
9GPT-5.4 Pro81.7%estimated ± 7.9 pp, low confidence
10Claude Mythos 581.7%estimated ± 7.9 pp, low confidence
11GPT-5.6 Terra81.7%estimated ± 7.9 pp, low confidence
12Muse Spark 1.281.6%estimated ± 13.6 pp, low confidence
13Ornith-1.5-397B81.4%measured
14Grok 4.580.7%estimated ± 13.6 pp, low confidence
15Gemini 3.8 Flash80.2%estimated ± 13.6 pp, low confidence
16Claude Sonnet 579.5%estimated ± 7.9 pp, medium confidence
17K-EXAONE 2.077.7%measured
18Gemini 3.7 Flash77.6%estimated ± 13.6 pp, low confidence
19Ornith-1.0-397B77.1%measured
20Quasar 438B75.2%estimated ± 13.6 pp, low confidence
21Claude Opus 4.875.0%estimated ± 3.2 pp, medium confidence
22MiniMax M374.5%measured
23dots3-note Preview73.4%measured
24Claude Opus 4.773.3%estimated ± 3.2 pp, medium confidence
25GLM-5.273.3%estimated ± 3.2 pp, medium confidence
26Gemini 3.6 Flash73.3%estimated ± 13.6 pp, low confidence
27Claude Opus 573.1%estimated ± 5.7 pp, low confidence
28Ornith-1.5-35B-A3B72.5%measured
29Qwen3.6-27B72.4%measured
30Inkling-Small71.1%estimated ± 5.7 pp, medium confidence
31Beam70.8%estimated ± 5.7 pp, medium confidence
32Claude Opus 4.670.4%measured
33Kimi K2.7 Code69.9%estimated ± 5.7 pp, medium confidence
34Ornith-1.0-35B69.8%measured
35Muse Glimmer 30B69.7%estimated ± 5.7 pp, medium confidence
36Hy369.3%estimated ± 13.6 pp, low confidence
37Inkling69.3%estimated ± 5.7 pp, medium confidence
38DeepSeek V4 Pro 081369.1%estimated ± 5.7 pp, medium confidence
39Qwen3.6-35B-A3B68.7%measured
40GPT-5.6 Luna68.5%estimated ± 7.9 pp, medium confidence
41Claude Sonnet 4.667.8%measured
42DeepSeek V4 Flash 073167.6%estimated ± 5.7 pp, medium confidence
43Step 3.7 Flash67.1%measured
44Ornith-1.5-9B66.5%measured
45Claude Opus 4.7 (Adaptive)66.2%estimated ± 5.6 pp, low confidence
46Hy4 preview66.2%estimated ± 5.6 pp, low confidence
47Kimi K366.2%estimated ± 5.6 pp, low confidence
48MiMo-V2.6-Flash66.2%estimated ± 5.6 pp, low confidence
49MiMo-V2.6-Pro66.2%estimated ± 5.6 pp, low confidence
50Muse Spark 1.166.2%estimated ± 5.6 pp, low confidence
51Muse Spark 1.366.2%estimated ± 5.6 pp, low confidence
52Qwen3.8-Flash-Next66.2%estimated ± 5.6 pp, low confidence
53Qwen3.8 Max66.2%estimated ± 5.6 pp, low confidence
54Step 5 Preview66.2%estimated ± 5.6 pp, low confidence
55GPT-5.266.0%estimated ± 5.6 pp, low confidence
56GPT-5.3 Codex65.5%estimated ± 5.6 pp, low confidence
57Qwen3.7 Max65.2%measured
58Solar Pro 465.1%estimated ± 5.7 pp, medium confidence
59Qwen3.8-27B65.0%estimated ± 5.6 pp, low confidence
60BTL-364.9%estimated ± 2.3 pp, low confidence
61Atria Dawn Preview64.7%estimated ± 2.3 pp, low confidence
62BTL-464.5%estimated ± 2.3 pp, medium confidence
63Ling 3.0 Flash64.5%estimated ± 2.3 pp, medium confidence
64Pokee-Isaac 28B64.4%estimated ± 2.3 pp, medium confidence
65LLaDA2.2-flash64.2%measured
66Ling 3.0 Flash FP864.1%estimated ± 13.6 pp, low confidence
67Solar Open 264.1%estimated ± 5.7 pp, medium confidence
68MiniCPM5-2B64.0%estimated ± 2.3 pp, medium confidence
69GPT-5.4 mini63.9%estimated ± 5.7 pp, medium confidence
70MiMo-V2.5-Pro63.8%measured
71Muse Spark63.8%measured
72Gemini 3.5 Flash63.7%estimated ± 3.2 pp, high confidence
73GPT-5.4 nano63.4%estimated ± 5.7 pp, medium confidence
74Granite 4.2 30B63.3%estimated ± 2.3 pp, medium confidence
75Ornith-1.0-9B63.1%measured
76LFM2.5-2.6B62.8%measured
77Agents-A1-4B62.8%estimated ± 7.1 pp, low confidence
78Agents-A162.7%estimated ± 7.1 pp, low confidence
79Qwen3.7 Plus62.7%measured
80Nemotron 3.5 Lightning 30B A3B NVFP462.4%estimated ± 7.9 pp, low confidence
81Nemotron 3 Ultra62.4%estimated ± 7.9 pp, low confidence
82Qwen3.5-122B-A10B62.4%estimated ± 7.9 pp, medium confidence
83GLM-5.162.3%measured
84Kimi K2.662.3%measured
85MiMo-V2.562.3%measured
86GPT-5.561.0%estimated ± 3.2 pp, high confidence
87Granite 4.2 3B60.6%estimated ± 2.3 pp, medium confidence
88Granite 4.2 8B60.6%estimated ± 2.3 pp, medium confidence
89GPT-5.460.3%measured
90LongCat-Flash-Lite-Sparse60.0%estimated ± 5.7 pp, medium confidence
91Claude Opus 4.559.6%measured
92Grok Build 0.159.4%estimated ± 6.8 pp, low confidence
93LFM2.5-8B-A1B59.1%estimated ± 2.3 pp, medium confidence
94Qwen3.6 Plus58.8%measured
95Grok 4.1 Fast58.7%estimated ± 6.8 pp, low confidence
96Gemini 3.1 Pro57.8%measured
97MiMo-V2-Pro57.8%measured
98GLM-557.7%measured
99LLaDA2.2-mini57.2%measured
100Qwen3 Max57.1%estimated ± 6.8 pp, low confidence
101Qwen3.5 397B56.8%measured
102Gemini 3.5 Flash-Lite56.8%estimated ± 13.6 pp, low confidence
103Grok 456.5%estimated ± 6.8 pp, low confidence
104Gemini 2.5 Pro56.3%estimated ± 6.8 pp, low confidence
105GPT-5.155.9%estimated ± 6.8 pp, low confidence
106GLM-5-Turbo55.8%measured
107Mellum2-12B-A2.5B-Thinking55.6%estimated ± 2.3 pp, low confidence
108Grok 4.155.4%estimated ± 3.2 pp, high confidence
109GLM-4.755.3%estimated ± 6.8 pp, low confidence
110Qwen3.5-27B55.0%estimated ± 6.8 pp, low confidence
111Mistral Medium 3.5 128B54.8%estimated ± 6.8 pp, low confidence
112Grok 4.354.5%estimated ± 3.2 pp, medium confidence
113Gemini 3.1 Flash-Lite54.5%estimated ± 6.8 pp, low confidence
114Grok 4.2054.5%estimated ± 6.8 pp, low confidence
115Mellum2-12B-A2.5B-Instruct53.9%estimated ± 2.3 pp, low confidence
116GLM-5V-Turbo53.8%measured
117Hy3 Preview53.6%estimated ± 6.8 pp, low confidence
118Gemma 4 31B52.7%estimated ± 6.8 pp, low confidence
119Kimi K2.552.3%measured
120Kimi K2.5 (Reasoning)51.0%estimated ± 6.8 pp, low confidence
121Trinity-Large-Thinking51.0%estimated ± 6.8 pp, low confidence
122Claude Sonnet 4.550.8%estimated ± 5.6 pp, low confidence
123GPT-5.1-Codex50.7%estimated ± 5.6 pp, low confidence
124GPT-5.2-Codex50.7%estimated ± 5.6 pp, low confidence
125Claude 4.1 Opus50.7%estimated ± 5.6 pp, low confidence
126Claude 4 Sonnet50.7%estimated ± 5.6 pp, low confidence
127Claude Haiku 4.550.7%estimated ± 5.6 pp, low confidence
128Gemini 3 Pro50.7%estimated ± 5.6 pp, low confidence
129GPT-5 (high)50.7%estimated ± 5.6 pp, low confidence
130Qwen3.5 Plus50.7%estimated ± 5.6 pp, low confidence
131Gemini 3 Flash49.2%measured
132GPT-OSS 120B48.9%estimated ± 6.8 pp, low confidence
133MiniMax M2.748.7%measured
134Qwen3.5-35B-A3B48.5%estimated ± 6.8 pp, low confidence
135GPT-4.145.8%estimated ± 6.8 pp, low confidence
136ZAYA1-8B45.7%estimated ± 2.3 pp, low confidence
137MiMo-V2-Omni45.2%measured
138DeepSeek V3.240.2%measured
139LFM2.5-VL-3B28.1%estimated ± 2.3 pp, low confidence
140Command A+21.7%estimated ± 13.6 pp, low confidence
141Mistral Large 315.6%estimated ± 13.6 pp, low confidence
142Mistral Small 49.8%estimated ± 13.6 pp, low confidence
143Mistral Small 4 (Reasoning)9.8%estimated ± 13.6 pp, low confidence
144GPT-OSS 20B9.5%estimated ± 13.6 pp, low confidence
145MiniCPM5-1B9.1%estimated ± 2.3 pp, low confidence
146Trinity-Large-Preview8.0%estimated ± 13.6 pp, low confidence
147Nemotron 3 Nano 30B7.0%estimated ± 13.6 pp, low confidence
148DeepSeek V35.6%estimated ± 13.6 pp, low confidence
149Nemotron 3 Super 100B5.5%measured
150Celeris-14.6%estimated ± 13.6 pp, low confidence
151Llama 4 Maverick4.4%estimated ± 13.6 pp, low confidence
152Llama 4 Scout4.0%estimated ± 13.6 pp, low confidence
153LFM2.5-VL-450M3.5%estimated ± 2.3 pp, low confidence
154LFM2.5-230M3.4%estimated ± 2.3 pp, low confidence
155Gemma 3 27B1.0%estimated ± 13.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General