benchgap
Coding

SWE-bench Pro leaderboard

As of 2026-10-07, the highest measured score on SWE-bench Pro is 89.9% by Claude Opus 5.5. 129 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.589.9%measured
2GPT-6 Astra86.7%estimated ± 5.1 pp, low confidence
3Gemini 4 Argon81.5%estimated ± 5.1 pp, medium confidence
4Claude Sonnet 5.581.3%measured
5Claude Fable 5.181.2%measured
6Claude Mythos 580.3%measured
7Claude Fable 580.0%measured
8Claude Opus 579.2%measured
9Pareto 26.977.1%estimated ± 7.8 pp, low confidence
10SWE-276.6%estimated ± 7.5 pp, medium confidence
11MiMo-V2.6-Pro76.1%estimated ± 6.8 pp, medium confidence
12Ember-174.7%estimated ± 3.2 pp, high confidence
13Grok 4.774.5%estimated ± 4.3 pp, high confidence
14Gemini 3.7 Flash74.1%estimated ± 4.8 pp, high confidence
15Sakana Fugu-Ultra73.7%measured
16Step 5 Preview73.0%estimated ± 6.8 pp, medium confidence
17Pareto 26.10 Preview73.0%estimated ± 7.8 pp, low confidence
18GPT-6 Sol71.0%estimated ± 6.8 pp, medium confidence
19Muse Spark 1.369.6%estimated ± 4.3 pp, high confidence
20Grok 4.669.4%estimated ± 4.3 pp, high confidence
21Claude Opus 4.869.2%measured
22SWE-1.768.8%estimated ± 7.5 pp, medium confidence
23Qwen3.8 Max67.7%measured
24Gemini 3.8 Flash67.4%estimated ± 4.3 pp, high confidence
25Claude Haiku 5.567.0%estimated ± 6.8 pp, medium confidence
26GPT-6 Luna66.3%estimated ± 6.8 pp, medium confidence
27GPT-6.1 Sol65.7%estimated ± 6.8 pp, medium confidence
28Mistral Large 465.7%estimated ± 6.8 pp, medium confidence
29Hy4 preview65.7%measured
30Ling 3.1 Flash65.6%estimated ± 6.8 pp, medium confidence
31Beam65.5%measured
32Ornith-1.5-397B65.1%measured
33Grok 4.564.7%measured
34GPT-5.6 Sol64.6%measured
35Qwen3.8 Max Preview64.4%estimated ± 6.0 pp, medium confidence
36Claude Opus 4.7 (Adaptive)64.3%measured
37GPT-5.6 Terra63.4%measured
38Qwen3.8-Omni-Flash63.3%measured
39Kimi K363.3%estimated ± 4.5 pp, high confidence
40Claude Sonnet 563.2%measured
41DeepSeek V4.1 Flash62.8%estimated ± 3.6 pp, medium confidence
42GPT-5.6 Luna62.7%measured
43Qwen3.8-Flash-Next62.5%measured
44GLM-5.362.4%estimated ± 3.6 pp, high confidence
45GLM-5.3-Flash62.2%estimated ± 3.6 pp, high confidence
46Ornith-1.0-397B62.2%measured
47GLM-5.262.1%measured
48Qwen3.8-27B61.7%measured
49Muse Spark 1.161.5%measured
50MiMo-V2.6-Flash61.2%estimated ± 6.8 pp, medium confidence
51dots3-note Preview61.0%measured
52Qwen3.7 Max60.6%measured
53Seed 2.1 Pro60.4%estimated ± 3.6 pp, high confidence
54Muse Spark 1.260.2%estimated ± 5.1 pp, medium confidence
55Atria Dawn Preview59.6%measured
56Ornith-1.5-35B-A3B59.6%measured
57Laguna S 2.159.4%measured
58Seed 2.1 Turbo59.0%estimated ± 3.6 pp, high confidence
59MiniMax M359.0%measured
60Sakana Fugu59.0%measured
61Quasar 438B58.9%estimated ± 6.0 pp, medium confidence
62Apodex 1.1 Mini58.7%estimated ± 6.0 pp, medium confidence
63GPT-5.558.6%measured
64Kimi K2.658.6%measured
65GLM-5.158.4%measured
66BTL-358.3%estimated ± 4.7 pp, high confidence
67GPT-5.457.7%measured
68Qwen3.7 Plus57.6%measured
69Hy357.6%estimated ± 6.0 pp, medium confidence
70Qwen 3.6 Max (preview)57.3%measured
71Claude Opus 4.757.2%estimated ± 2.1 pp, high confidence
72MiMo-V2.5-Pro57.2%measured
73Claude Sonnet 4.657.1%estimated ± 2.1 pp, high confidence
74Claude Opus 4.557.1%measured
75GPT-5.4 mini57.1%estimated ± 2.1 pp, high confidence
76Gemini 3.6 Flash57.0%estimated ± 4.5 pp, high confidence
77Grok 4.357.0%estimated ± 4.8 pp, high confidence
78Qwen3.5 Flash57.0%estimated ± 4.8 pp, high confidence
79GLM-4.557.0%estimated ± 4.8 pp, medium confidence
80GPT-5.2-Codex56.9%estimated ± 2.1 pp, high confidence
81Gemini 3.1 Pro56.8%estimated ± 2.1 pp, high confidence
82GPT-5.3 Codex56.8%measured
83GPT-5.4 nano56.6%estimated ± 2.1 pp, high confidence
84Ling 3.0 Flash56.6%measured
85Qwen3.6 Plus56.6%measured
86GPT-5.1-Codex-Max56.5%estimated ± 2.1 pp, high confidence
87Gemini 3 Flash56.4%estimated ± 2.1 pp, high confidence
88Step 3.7 Flash56.3%measured
89MiniMax M2.756.2%measured
90MiMo-V2.556.1%measured
91Inkling-Small55.9%measured
92GPT-5.1-Codex55.7%estimated ± 2.1 pp, high confidence
93GPT-5.255.6%measured
94DeepSeek V4 Pro 081355.4%measured
95Gemini 3.5 Flash55.1%measured
96GLM-555.1%measured
97BTL-455.0%estimated ± 3.2 pp, high confidence
98Kimi K2.7 Code54.8%estimated ± 4.5 pp, high confidence
99MiMo-V2-Pro54.4%estimated ± 3.2 pp, high confidence
100Inkling54.3%measured
101Gemini 3.5 Flash-Lite54.2%measured
102Apodex 1.154.1%estimated ± 3.2 pp, high confidence
103Mistral Medium 3.5 128B53.9%estimated ± 3.2 pp, high confidence
104Qwen3.6-27B53.5%measured
105Claude Sonnet 4.553.4%estimated ± 3.2 pp, high confidence
106Claude Opus 4.653.4%measured
107Kimi K2.5 (Reasoning)52.9%estimated ± 3.2 pp, high confidence
108MAI-Thinking-152.8%measured
109Ling 3.0 Flash FP852.6%estimated ± 6.0 pp, medium confidence
110DeepSeek V4 Flash 073152.6%measured
111Composer 2.552.4%estimated ± 4.3 pp, medium confidence
112Muse Spark52.4%measured
113GPT-5.151.8%estimated ± 6.0 pp, medium confidence
114Grok 4.2051.8%measured
115Muse Glimmer 30B51.2%measured
116Qwen3.5 397B50.9%measured
117A.X K250.8%estimated ± 4.7 pp, high confidence
118Kimi K2.550.7%measured
119MiMo-V2-Omni50.4%estimated ± 3.2 pp, high confidence
120Ornith-1.0-35B50.4%measured
121GLM-4.650.3%estimated ± 2.1 pp, high confidence
122Ling 3.0 Flash VL50.2%estimated ± 6.8 pp, medium confidence
123Claude 4.1 Opus50.1%estimated ± 3.2 pp, high confidence
124Hy3 Preview50.0%estimated ± 3.2 pp, high confidence
125Qwen3.6-35B-A3B49.5%measured
126GLM-4.749.2%estimated ± 3.2 pp, high confidence
127Laguna M.149.2%measured
128MiMo-V2-Flash48.7%estimated ± 3.2 pp, high confidence
129Claude Haiku 4.548.6%estimated ± 3.2 pp, high confidence
130Claude 4 Sonnet47.9%estimated ± 3.2 pp, high confidence
131MAI-Code-1.1-Flash47.8%estimated ± 3.2 pp, high confidence
132Gemma 4 31B47.7%estimated ± 6.0 pp, medium confidence
133Laguna XS 2.147.6%measured
134Qwen3.5-27B47.6%estimated ± 3.2 pp, high confidence
135Ornith-1.5-9B47.5%measured
136Qwen3.5-122B-A10B47.1%estimated ± 3.2 pp, high confidence
137Nemotron 3 Ultra47.0%estimated ± 3.2 pp, high confidence
138Laguna XS.246.3%measured
139Grok Code Fast 145.7%estimated ± 3.2 pp, high confidence
140Solar Pro 445.5%estimated ± 3.2 pp, high confidence
141Solar Open 245.2%estimated ± 3.2 pp, high confidence
142o144.9%estimated ± 6.0 pp, medium confidence
143Gemma 4 26B A4B44.6%estimated ± 6.0 pp, medium confidence
144Qwen3.5-35B-A3B43.9%estimated ± 3.2 pp, high confidence
145GPT-5 (high)43.4%estimated ± 6.0 pp, medium confidence
146Nemotron 3 Super 100B43.4%estimated ± 6.0 pp, medium confidence
147Ornith-1.0-9B42.9%measured
148K-EXAONE 2.042.7%estimated ± 3.2 pp, high confidence
149DeepSeek V3 032442.2%estimated ± 6.8 pp, medium confidence
150North Mini Code41.9%estimated ± 6.8 pp, medium confidence
151Mercury 2.541.4%estimated ± 6.8 pp, medium confidence
152LongCat-Flash-Lite-Sparse40.6%measured
153o1-preview40.4%estimated ± 6.0 pp, medium confidence
154K-Exaone38.7%estimated ± 6.0 pp, medium confidence
155Gemini 2.5 Pro38.0%estimated ± 3.2 pp, high confidence
156GPT-OSS 120B37.3%estimated ± 6.0 pp, medium confidence
157Command A+34.9%estimated ± 6.0 pp, medium confidence
158Ternary Bonsai 2 27B34.9%estimated ± 3.2 pp, high confidence
159Mistral Small 433.7%estimated ± 6.0 pp, medium confidence
160Mistral Small 4 (Reasoning)33.7%estimated ± 6.0 pp, medium confidence
161Granite 4.2 30B33.3%measured
162Trinity-Large-Preview32.9%estimated ± 6.0 pp, medium confidence
163Trinity-Large-Thinking32.9%estimated ± 6.0 pp, medium confidence
164Ling 2.6 Flash32.4%estimated ± 6.0 pp, medium confidence
165Gemini 1.5 Pro30.8%estimated ± 6.0 pp, medium confidence
166LLaDA2.2-flash30.1%measured
167GPT-4.128.8%estimated ± 3.2 pp, high confidence
168GPT-4 Turbo28.6%estimated ± 6.0 pp, low confidence
169GPT-OSS 20B27.7%estimated ± 6.0 pp, low confidence
170ZAYA1-74B-Preview27.5%estimated ± 3.2 pp, high confidence
171Nemotron 3.5 Lightning 30B A3B NVFP427.2%estimated ± 3.2 pp, high confidence
172Mistral Large 327.0%estimated ± 6.0 pp, low confidence
173Claude 3 Opus26.4%estimated ± 6.0 pp, low confidence
174o3-mini24.0%estimated ± 3.2 pp, high confidence
175Claude 3.5 Sonnet23.8%estimated ± 3.2 pp, high confidence
176Llama 4 Maverick22.8%estimated ± 6.0 pp, low confidence
177Solar Pro 321.3%estimated ± 6.8 pp, low confidence
178Celeris-120.5%estimated ± 6.0 pp, low confidence
179Nemotron 3 Nano 30B20.5%estimated ± 6.0 pp, low confidence
180Gemma 4 12B20.1%estimated ± 4.7 pp, high confidence
181Nemotron 3 Nano Omni 30B A3B19.7%estimated ± 6.0 pp, low confidence
182Ling 3.0 Tiny19.3%estimated ± 6.8 pp, low confidence
183Granite 4.2 8B19.1%measured
184DeepSeek V317.9%estimated ± 3.2 pp, medium confidence
185Ultravox v0.6 Llama 3.3 70B17.4%estimated ± 6.0 pp, low confidence
186GPT-4o mini16.7%estimated ± 6.0 pp, low confidence
187GPT-4.1 nano16.4%estimated ± 6.0 pp, low confidence
188Mellum2-12B-A2.5B-Thinking15.8%estimated ± 4.7 pp, high confidence
189Granite 4.2 3B15.4%estimated ± 4.7 pp, high confidence
190Gemma 3 27B15.0%estimated ± 6.0 pp, low confidence
191MiniCPM5-2B14.4%measured
192Gemma 4 E4B14.1%estimated ± 6.0 pp, low confidence
193Llama 4 Scout12.4%estimated ± 6.0 pp, low confidence
194Gemma 4 E2B11.1%estimated ± 6.0 pp, low confidence
195ZAYA1-8B9.3%estimated ± 4.7 pp, medium confidence
196GPT-4.1 mini4.5%estimated ± 3.2 pp, medium confidence
197Agents-A1-4B3.8%estimated ± 4.7 pp, medium confidence
198LFM2.5-2.6B3.7%estimated ± 4.7 pp, medium confidence
199Mellum2-12B-A2.5B-Instruct0.1%estimated ± 4.7 pp, medium confidence
200MiniCPM5-1B0.1%estimated ± 4.7 pp, medium confidence
201LLaDA2.2-mini0.0%estimated ± 4.7 pp, medium confidence
202Gemini 3.1 Flash-Lite0.0%estimated ± 2.1 pp, high confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General