benchgap
Coding

SWE-bench Verified leaderboard

As of 2026-10-07, the highest measured score on SWE-bench Verified is 96.0% by Claude Opus 5. 140 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.596.5%estimated ± 0.8 pp, low confidence
2Claude Opus 596.0%measured
3Gemini 3.8 Flash95.8%estimated ± 4.4 pp, high confidence
4Claude Mythos 595.5%measured
5GPT-6 Astra95.3%estimated ± 0.8 pp, medium confidence
6Claude Fable 595.0%measured
7Claude Fable 5.194.6%estimated ± 2.5 pp, medium confidence
8Muse Spark 1.394.3%estimated ± 5.2 pp, low confidence
9Pareto 26.993.4%estimated ± 5.2 pp, low confidence
10Ember-192.2%measured
11GPT-6.1 Sol91.9%estimated ± 5.2 pp, low confidence
12MiMo-V2.6-Pro91.9%estimated ± 5.2 pp, low confidence
13SWE-291.8%estimated ± 0.8 pp, medium confidence
14Grok 4.791.3%estimated ± 5.2 pp, low confidence
15Sakana Fugu-Ultra91.1%estimated ± 2.5 pp, high confidence
16Pareto 26.10 Preview90.6%estimated ± 5.2 pp, low confidence
17Grok 4.690.5%estimated ± 4.4 pp, high confidence
18GPT-6 Sol89.8%estimated ± 5.2 pp, low confidence
19Gemini 3.6 Flash89.6%estimated ± 4.4 pp, high confidence
20MiMo-V2.6-Flash89.2%estimated ± 5.2 pp, low confidence
21Step 5 Preview89.1%estimated ± 5.2 pp, low confidence
22Gemini 4 Argon89.0%estimated ± 4.4 pp, medium confidence
23Claude Opus 4.888.6%measured
24Claude Haiku 5.588.4%estimated ± 0.8 pp, medium confidence
25GPT-6 Luna88.4%estimated ± 5.2 pp, low confidence
26Claude Sonnet 5.588.3%estimated ± 0.8 pp, medium confidence
27Claude Opus 4.7 (Adaptive)87.6%measured
28Ling 3.1 Flash87.2%estimated ± 6.6 pp, medium confidence
29Mistral Large 486.4%estimated ± 4.4 pp, medium confidence
30Gemini 3.7 Flash86.2%estimated ± 0.8 pp, medium confidence
31Ornith-1.5-397B86.0%measured
32Grok 4.585.8%estimated ± 2.5 pp, high confidence
33GPT-5.585.7%estimated ± 0.8 pp, medium confidence
34GPT-5.6 Sol85.7%estimated ± 2.5 pp, high confidence
35SWE-1.785.3%estimated ± 0.8 pp, medium confidence
36Qwen3.8 Max Preview85.2%estimated ± 5.5 pp, medium confidence
37Claude Sonnet 585.2%measured
38GPT-5.3 Codex85.0%measured
39Muse Spark 1.284.9%estimated ± 4.5 pp, high confidence
40GPT-5.6 Terra84.9%estimated ± 2.5 pp, high confidence
41GPT-5.6 Luna84.4%estimated ± 2.5 pp, high confidence
42DeepSeek V4.1 Flash84.1%estimated ± 2.4 pp, medium confidence
43Muse Spark 1.183.5%estimated ± 2.5 pp, high confidence
44Claude Opus 4.783.1%estimated ± 0.8 pp, medium confidence
45Hy4 preview82.9%estimated ± 2.4 pp, high confidence
46GLM-5.382.7%estimated ± 2.4 pp, high confidence
47Ornith-1.0-397B82.4%measured
48GLM-5.3-Flash82.3%estimated ± 2.4 pp, high confidence
49Qwen3.8 Max82.2%estimated ± 2.4 pp, high confidence
50Atria Dawn Preview82.2%estimated ± 2.5 pp, high confidence
51Laguna S 2.182.0%estimated ± 2.5 pp, high confidence
52Quasar 438B81.9%estimated ± 5.5 pp, medium confidence
53Apodex 1.1 Mini81.8%estimated ± 5.5 pp, medium confidence
54Claude Opus 4.6 (Adaptive)81.7%estimated ± 4.4 pp, high confidence
55Sakana Fugu81.7%estimated ± 2.5 pp, high confidence
56Composer 2.581.4%estimated ± 4.5 pp, high confidence
57Kimi K381.2%estimated ± 4.4 pp, high confidence
58Hy381.0%estimated ± 5.5 pp, medium confidence
59Claude Opus 4.580.9%measured
60Beam80.9%measured
61Claude Opus 4.680.8%measured
62GPT-5.480.7%estimated ± 2.5 pp, high confidence
63DeepSeek V4 Pro 081380.6%measured
64MiniMax M380.5%measured
65GLM-5.280.4%estimated ± 2.4 pp, high confidence
66Qwen3.8-Omni-Flash80.4%estimated ± 2.4 pp, high confidence
67Qwen3.7 Max80.4%measured
68MiMo-V2.5-Pro80.3%estimated ± 2.5 pp, high confidence
69GPT-5.4 mini80.2%estimated ± 0.8 pp, medium confidence
70Inkling-Small80.2%measured
71Kimi K2.680.2%measured
72Qwen3.8-Flash-Next80.2%estimated ± 2.4 pp, high confidence
73GPT-5.280.0%measured
74Seed 2.1 Pro79.9%estimated ± 2.4 pp, high confidence
75Ling 3.0 Flash79.8%estimated ± 2.5 pp, high confidence
76Claude Sonnet 4.679.6%measured
77Step 3.7 Flash79.6%estimated ± 2.5 pp, high confidence
78MiMo-V2.579.4%estimated ± 2.5 pp, high confidence
79DeepSeek V4 Flash 073179.0%measured
80Ornith-1.5-35B-A3B79.0%measured
81Seed 2.1 Turbo78.8%estimated ± 2.4 pp, high confidence
82GPT-5.2-Codex78.8%estimated ± 4.4 pp, high confidence
83Qwen3.6 Plus78.8%measured
84Gemini 3.5 Flash78.6%estimated ± 2.5 pp, high confidence
85Qwen 3.6 Max (preview)78.5%estimated ± 2.4 pp, high confidence
86GLM-5.178.5%estimated ± 2.4 pp, high confidence
87BTL-478.4%measured
88dots3-note Preview78.4%measured
89Qwen3.8-27B78.3%estimated ± 2.4 pp, high confidence
90DeepSeek V3.278.0%estimated ± 4.8 pp, medium confidence
91MiMo-V2-Pro78.0%measured
92Gemini 3.5 Flash-Lite77.8%estimated ± 2.5 pp, high confidence
93GLM-577.8%measured
94Apodex 1.177.7%measured
95Qwen3.7 Plus77.7%measured
96Gemini 3.1 Pro77.7%estimated ± 4.4 pp, high confidence
97Inkling77.6%measured
98Mistral Medium 3.5 128B77.6%measured
99Muse Spark77.4%measured
100MiniMax M2.777.4%estimated ± 2.4 pp, high confidence
101Claude Sonnet 4.577.2%measured
102Qwen3.6-27B77.2%measured
103Ling 3.0 Flash FP876.8%estimated ± 5.5 pp, medium confidence
104Kimi K2.576.8%measured
105Kimi K2.5 (Reasoning)76.8%measured
106Grok 4.2076.7%measured
107GPT-5.4 nano76.6%estimated ± 4.4 pp, high confidence
108GPT-5.176.3%estimated ± 4.4 pp, high confidence
109Grok 4.376.3%estimated ± 4.4 pp, high confidence
110Qwen3.5 Flash76.3%estimated ± 4.4 pp, high confidence
111Kimi K2.7 Code76.3%estimated ± 4.4 pp, high confidence
112GLM-4.576.3%estimated ± 4.4 pp, high confidence
113Qwen3.5 397B76.2%measured
114GLM-5 (Reasoning)76.1%estimated ± 4.4 pp, high confidence
115Muse Glimmer 30B76.0%measured
116Claude Sonnet 4.5 Thinking75.9%estimated ± 4.4 pp, high confidence
117GPT-5.1-Codex-Max75.8%estimated ± 4.4 pp, high confidence
118Composer 275.7%estimated ± 4.8 pp, medium confidence
119Ornith-1.0-35B75.6%measured
120Claude Opus 4.5 Thinking75.5%estimated ± 4.4 pp, high confidence
121Gemini 3 Flash75.5%estimated ± 4.4 pp, high confidence
122GPT-5 (high)75.4%estimated ± 4.4 pp, high confidence
123MiMo-V2-Omni74.8%measured
124Qwen3.5 Plus74.6%estimated ± 4.4 pp, high confidence
125Laguna M.174.6%measured
126Claude 4.1 Opus74.5%measured
127MiniMax M2.574.5%estimated ± 4.4 pp, high confidence
128Hy3 Preview74.4%measured
129Gemini 3 Pro74.4%estimated ± 4.4 pp, high confidence
130GPT-5 mini74.3%estimated ± 4.4 pp, high confidence
131GPT-5.1-Codex74.1%estimated ± 4.4 pp, high confidence
132Claude Haiku 4.5 Thinking73.8%estimated ± 4.4 pp, high confidence
133GLM-4.773.8%measured
134MAI-Thinking-173.5%measured
135MiMo-V2-Flash73.4%measured
136Qwen3.6-35B-A3B73.4%measured
137Claude Haiku 4.573.3%measured
138Claude 4 Sonnet72.7%measured
139DeepSeek V3.2 (Thinking)72.6%estimated ± 4.4 pp, high confidence
140MAI-Code-1.1-Flash72.6%measured
141Qwen3.5-27B72.4%measured
142Ling 3.0 Flash VL72.4%estimated ± 6.6 pp, medium confidence
143Qwen3 Max72.3%estimated ± 4.4 pp, high confidence
144GLM-4.672.2%estimated ± 4.4 pp, high confidence
145BTL-372.2%estimated ± 10.5 pp, low confidence
146Qwen3.5-122B-A10B72.0%measured
147Nemotron 3 Ultra71.9%measured
148Grok 4.1 Fast (Reasoning)71.9%estimated ± 4.4 pp, high confidence
149Gemini 3.1 Flash-Lite71.7%estimated ± 4.4 pp, medium confidence
150Grok 4 Fast (Reasoning)71.7%estimated ± 4.4 pp, medium confidence
151Laguna XS 2.170.9%measured
152Grok Code Fast 170.8%measured
153Ornith-1.5-9B70.6%measured
154Solar Pro 470.6%measured
155Solar Open 270.4%measured
156Laguna XS.269.9%measured
157Ornith-1.0-9B69.4%measured
158Qwen3.5-35B-A3B69.2%measured
159K-EXAONE 2.068.2%measured
160LongCat-Flash-Lite-Sparse68.2%measured
161o168.1%estimated ± 5.5 pp, medium confidence
162Gemma 4 26B A4B67.7%estimated ± 5.5 pp, medium confidence
163A.X K266.4%estimated ± 5.7 pp, medium confidence
164Nemotron 3 Super 100B65.9%estimated ± 5.5 pp, medium confidence
165DeepSeek V3 032464.6%estimated ± 6.6 pp, medium confidence
166North Mini Code64.3%estimated ± 6.6 pp, medium confidence
167Gemini 2.5 Pro63.8%measured
168Mercury 2.562.2%estimated ± 5.7 pp, medium confidence
169Mellum2-12B-A2.5B-Thinking61.9%estimated ± 10.5 pp, low confidence
170o1-preview61.3%estimated ± 5.5 pp, medium confidence
171Gemma 4 31B60.8%estimated ± 4.8 pp, low confidence
172Ternary Bonsai 2 27B60.8%measured
173ZAYA1-8B59.4%estimated ± 10.5 pp, low confidence
174K-Exaone58.5%estimated ± 5.5 pp, medium confidence
175Granite 4.2 30B57.0%measured
176Gemma 4 12B56.8%estimated ± 5.5 pp, medium confidence
177GPT-OSS 120B55.9%estimated ± 5.5 pp, medium confidence
178GPT-4.154.6%measured
179ZAYA1-74B-Preview53.2%measured
180Nemotron 3.5 Lightning 30B A3B NVFP452.8%measured
181Command A+51.5%estimated ± 5.5 pp, medium confidence
182Agents-A1-4B50.4%estimated ± 5.7 pp, medium confidence
183o3-mini49.3%measured
184LLaDA2.2-flash49.3%measured
185Mistral Small 449.2%estimated ± 5.5 pp, medium confidence
186Mistral Small 4 (Reasoning)49.2%estimated ± 5.5 pp, medium confidence
187Claude 3.5 Sonnet49.0%measured
188Granite 4.2 8B47.7%measured
189Trinity-Large-Preview47.5%estimated ± 5.5 pp, medium confidence
190Trinity-Large-Thinking47.5%estimated ± 5.5 pp, medium confidence
191Ling 2.6 Flash46.5%estimated ± 5.5 pp, medium confidence
192MiniCPM5-2B46.4%measured
193Solar Pro 344.4%estimated ± 6.6 pp, low confidence
194Gemini 1.5 Pro43.1%estimated ± 5.5 pp, medium confidence
195Granite 4.2 3B42.7%estimated ± 5.7 pp, low confidence
196Ling 3.0 Tiny42.4%estimated ± 6.6 pp, low confidence
197DeepSeek V342.0%measured
198Mellum2-12B-A2.5B-Instruct38.5%estimated ± 10.5 pp, low confidence
199GPT-4 Turbo38.3%estimated ± 5.5 pp, medium confidence
200GPT-OSS 20B36.5%estimated ± 5.5 pp, medium confidence
201MiniCPM5-1B35.4%estimated ± 10.5 pp, low confidence
202Mistral Large 335.0%estimated ± 5.5 pp, low confidence
203Claude 3 Opus33.7%estimated ± 5.5 pp, low confidence
204LLaDA2.2-mini30.6%estimated ± 10.5 pp, low confidence
205Llama 4 Maverick25.7%estimated ± 5.5 pp, low confidence
206GPT-4.1 mini23.6%measured
207Celeris-121.0%estimated ± 5.5 pp, low confidence
208Nemotron 3 Nano 30B21.0%estimated ± 5.5 pp, low confidence
209Nemotron 3 Nano Omni 30B A3B19.5%estimated ± 5.5 pp, low confidence
210Ultravox v0.6 Llama 3.3 70B15.1%estimated ± 5.5 pp, low confidence
211GPT-4o mini13.9%estimated ± 5.5 pp, low confidence
212GPT-4.1 nano13.3%estimated ± 5.5 pp, low confidence
213Gemma 3 27B11.0%estimated ± 5.5 pp, low confidence
214Gemma 4 E4B9.6%estimated ± 5.5 pp, low confidence
215Llama 4 Scout7.3%estimated ± 5.5 pp, low confidence
216LFM2.5-2.6B6.6%estimated ± 5.5 pp, low confidence
217Gemma 4 E2B5.7%estimated ± 5.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General