benchgap
Agentic · tools

CWE-bench v1 leaderboard

As of 2026-10-07, the highest measured score on CWE-bench v1 is 68.0% by Gemini 4 Argon. 114 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 4 Argon68.0%measured
2GPT-6 Astra68.0%measured
3Grok 4.768.0%measured
4Claude Opus 5.567.0%measured
5Claude Sonnet 5.565.0%estimated ± 6.0 pp, medium confidence
6Claude Opus 562.7%estimated ± 6.9 pp, medium confidence
7Qwen3.8 Max Preview61.2%estimated ± 6.9 pp, medium confidence
8GPT-6.1 Sol60.2%estimated ± 6.0 pp, medium confidence
9Qwen3.8-Flash-Next60.0%estimated ± 6.9 pp, medium confidence
10Ling 3.1 Flash59.7%estimated ± 6.9 pp, medium confidence
11Claude Fable 559.4%estimated ± 6.9 pp, medium confidence
12GPT-5.6 Sol59.4%estimated ± 6.9 pp, medium confidence
13MiMo-V2.6-Flash59.4%estimated ± 6.9 pp, medium confidence
14DeepSeek V4 Pro 081358.8%estimated ± 6.9 pp, medium confidence
15GLM-5.358.1%estimated ± 6.0 pp, medium confidence
16Claude Fable 5.158.0%measured
17Grok 4.657.0%measured
18GLM-5.3-Flash56.6%estimated ± 6.0 pp, medium confidence
19Gemini 3.8 Flash56.2%estimated ± 6.0 pp, medium confidence
20Mistral Large 456.2%estimated ± 6.0 pp, medium confidence
21MiMo-V2.6-Pro55.1%estimated ± 6.0 pp, medium confidence
22Muse Spark 1.255.1%estimated ± 6.9 pp, medium confidence
23DeepSeek V4.1 Flash55.0%measured
24Muse Spark 1.355.0%measured
25Kimi K354.9%estimated ± 6.0 pp, medium confidence
26Claude Sonnet 554.7%estimated ± 6.9 pp, medium confidence
27GPT-5.6 Luna54.6%estimated ± 6.9 pp, medium confidence
28Claude Opus 4.854.3%estimated ± 6.9 pp, medium confidence
29GPT-5.6 Terra54.3%estimated ± 6.9 pp, medium confidence
30DeepSeek V4 Flash 073153.7%estimated ± 6.9 pp, medium confidence
31Hy4 preview53.0%measured
32Gemini 3.7 Flash52.1%estimated ± 6.9 pp, medium confidence
33GPT-6 Sol52.0%measured
34Grok 4.552.0%estimated ± 6.9 pp, medium confidence
35GLM-5.251.4%estimated ± 6.9 pp, medium confidence
36Claude Opus 4.7 (Adaptive)50.7%estimated ± 6.9 pp, medium confidence
37GPT-5.550.7%estimated ± 6.9 pp, medium confidence
38GPT-6 Luna50.5%estimated ± 6.0 pp, medium confidence
39Gemini 3.5 Flash50.3%estimated ± 6.9 pp, medium confidence
40Step 5 Preview48.7%estimated ± 6.0 pp, medium confidence
41Gemini 3.6 Flash48.0%estimated ± 6.9 pp, medium confidence
42Qwen3.8-27B46.6%estimated ± 6.0 pp, medium confidence
43GPT-5.446.5%estimated ± 6.9 pp, medium confidence
44Hy3 Preview45.1%estimated ± 6.9 pp, medium confidence
45Muse Spark 1.145.0%estimated ± 6.9 pp, medium confidence
46Apodex 1.144.3%estimated ± 6.9 pp, medium confidence
47Apodex 1.1 Mini44.3%estimated ± 6.9 pp, medium confidence
48Quasar 438B43.7%estimated ± 6.9 pp, medium confidence
49Ling 3.0 Flash VL42.9%estimated ± 6.9 pp, medium confidence
50Qwen3.7 Max41.4%estimated ± 6.9 pp, medium confidence
51Inkling-Small41.0%estimated ± 6.9 pp, medium confidence
52MiMo-V2.5-Pro41.0%estimated ± 6.9 pp, medium confidence
53GLM-5.140.8%estimated ± 6.9 pp, medium confidence
54Solar Pro 440.4%estimated ± 6.9 pp, medium confidence
55Claude Haiku 5.539.6%estimated ± 6.0 pp, medium confidence
56Grok 4.339.1%estimated ± 6.9 pp, medium confidence
57Hy338.3%estimated ± 6.9 pp, low confidence
58MiniMax M337.2%estimated ± 6.0 pp, medium confidence
59Muse Glimmer 30B37.0%estimated ± 6.0 pp, medium confidence
60Nemotron 3 Ultra37.0%estimated ± 6.0 pp, low confidence
61Inkling37.0%measured
62Kimi K2.637.0%estimated ± 6.9 pp, low confidence
63Kimi K2.7 Code37.0%estimated ± 6.9 pp, low confidence
64GLM-4.735.7%estimated ± 6.9 pp, low confidence
65GPT-5.4 mini35.7%estimated ± 6.9 pp, low confidence
66Step 3.7 Flash35.7%estimated ± 6.9 pp, low confidence
67MiniMax M2.735.6%estimated ± 6.9 pp, low confidence
68Muse Spark35.0%estimated ± 6.9 pp, low confidence
69Qwen3.6 Plus34.6%estimated ± 6.9 pp, low confidence
70Qwen3.6-27B34.3%estimated ± 6.9 pp, low confidence
71Gemini 3.5 Flash-Lite34.2%estimated ± 6.9 pp, low confidence
72A.X K233.0%estimated ± 6.9 pp, low confidence
73GPT-5.4 nano32.2%estimated ± 6.9 pp, low confidence
74Ling 3.0 Flash32.1%estimated ± 6.9 pp, low confidence
75Ling 3.0 Flash FP832.1%estimated ± 6.9 pp, low confidence
76GPT-5 (high)30.6%estimated ± 6.9 pp, low confidence
77Qwen3.6-35B-A3B29.2%estimated ± 6.9 pp, low confidence
78Kimi K2.526.1%estimated ± 6.9 pp, low confidence
79Kimi K2.5 (Reasoning)26.1%estimated ± 6.9 pp, low confidence
80GPT-5.125.2%estimated ± 6.9 pp, low confidence
81Qwen3.5-122B-A10B24.3%estimated ± 6.9 pp, low confidence
82Gemini 3.1 Pro22.9%estimated ± 6.9 pp, low confidence
83Qwen3.7 Plus21.3%estimated ± 6.9 pp, low confidence
84Mistral Medium 3.5 128B20.9%estimated ± 6.9 pp, low confidence
85MiniCPM5-2B17.2%estimated ± 6.9 pp, low confidence
86Nemotron 3.5 Lightning 30B A3B NVFP411.9%estimated ± 6.9 pp, low confidence
87MiMo-V2-Flash11.3%estimated ± 6.9 pp, low confidence
88Gemma 4 31B10.5%estimated ± 6.9 pp, low confidence
89GPT-OSS 120B9.7%estimated ± 6.9 pp, low confidence
90Granite 4.2 30B6.9%estimated ± 6.9 pp, low confidence
91Gemma 4 26B A4B6.1%estimated ± 6.9 pp, low confidence
92Ling 3.0 Tiny5.7%estimated ± 6.9 pp, low confidence
93Command A+1.7%estimated ± 6.9 pp, low confidence
94Celeris-10.0%estimated ± 6.9 pp, low confidence
95DeepSeek V30.0%estimated ± 6.9 pp, low confidence
96DeepSeek V3 03240.0%estimated ± 6.9 pp, low confidence
97Gemini 2.5 Pro0.0%estimated ± 6.9 pp, low confidence
98Gemma 3 27B0.0%estimated ± 6.9 pp, low confidence
99Gemma 4 12B0.0%estimated ± 6.9 pp, low confidence
100Gemma 4 E2B0.0%estimated ± 6.9 pp, low confidence
101Gemma 4 E4B0.0%estimated ± 6.9 pp, low confidence
102GPT-4.1 mini0.0%estimated ± 6.9 pp, low confidence
103GPT-4.1 nano0.0%estimated ± 6.9 pp, low confidence
104GPT-4o0.0%estimated ± 6.9 pp, low confidence
105GPT-4o mini0.0%estimated ± 6.9 pp, low confidence
106GPT-OSS 20B0.0%estimated ± 6.9 pp, low confidence
107Granite 4.2 3B0.0%estimated ± 6.9 pp, low confidence
108Granite 4.2 8B0.0%estimated ± 6.9 pp, low confidence
109K-Exaone0.0%estimated ± 6.9 pp, low confidence
110LFM2.5-2.6B0.0%estimated ± 6.9 pp, low confidence
111Ling 2.6 Flash0.0%estimated ± 6.9 pp, low confidence
112Llama 4 Maverick0.0%estimated ± 6.9 pp, low confidence
113Llama 4 Scout0.0%estimated ± 6.9 pp, low confidence
114Mercury 2.50.0%estimated ± 6.9 pp, low confidence
115Mistral Large 30.0%estimated ± 6.9 pp, low confidence
116Mistral Small 40.0%estimated ± 6.9 pp, low confidence
117Mistral Small 4 (Reasoning)0.0%estimated ± 6.9 pp, low confidence
118Nemotron 3 Nano 30B0.0%estimated ± 6.9 pp, low confidence
119Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 6.9 pp, low confidence
120Nemotron 3 Super 100B0.0%estimated ± 6.9 pp, low confidence
121North Mini Code0.0%estimated ± 6.9 pp, low confidence
122Solar Pro 30.0%estimated ± 6.9 pp, low confidence
123Trinity-Large-Preview0.0%estimated ± 6.9 pp, low confidence
124Trinity-Large-Thinking0.0%estimated ± 6.9 pp, low confidence
125Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 6.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General