benchgap
Coding

DeepSWE leaderboard

As of 2026-10-07, the highest measured score on DeepSWE is 77.9% by Gemini 4 Argon. 107 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 4 Argon77.9%measured
2Muse Spark 1.375.4%measured
3Ember-175.2%measured
4Claude Opus 5.574.2%measured
5DeepSeek V4.1 Flash74.2%measured
6GPT-6 Astra74.1%measured
7Pareto 26.974.0%measured
8Gemini 3.8 Flash73.8%measured
9SWE-273.0%measured
10Claude Fable 572.9%estimated ± 3.4 pp, medium confidence
11GPT-5.6 Sol72.7%measured
12GPT-6.1 Sol71.9%measured
13MiMo-V2.6-Pro71.9%measured
14Claude Sonnet 5.571.0%measured
15Grok 4.771.0%measured
16Pareto 26.10 Preview69.9%measured
17GPT-5.6 Terra69.6%measured
18Claude Opus 4.869.4%estimated ± 3.4 pp, medium confidence
19Claude Haiku 5.569.3%estimated ± 3.4 pp, medium confidence
20Claude Opus 568.8%measured
21GPT-6 Sol68.8%measured
22MiMo-V2.6-Flash67.9%measured
23Step 5 Preview67.7%measured
24Kimi K367.5%measured
25Claude Fable 5.167.4%measured
26GPT-5.6 Luna67.2%measured
27GLM-5.366.9%measured
28GPT-6 Luna66.6%measured
29GPT-5.566.0%estimated ± 3.4 pp, low confidence
30Grok 4.665.9%measured
31Claude Sonnet 565.7%estimated ± 3.4 pp, low confidence
32Gemini 3.7 Flash65.3%measured
33SWE-1.765.2%estimated ± 3.4 pp, low confidence
34GLM-5.264.9%estimated ± 4.4 pp, high confidence
35Hy4 preview64.3%measured
36GLM-5.3-Flash63.4%measured
37Claude Opus 4.7 (Adaptive)62.8%estimated ± 6.4 pp, medium confidence
38DeepSeek V4 Pro 081362.7%measured
39Muse Spark 1.259.3%measured
40Qwen3.8 Max Preview59.3%estimated ± 6.4 pp, medium confidence
41Claude Opus 4.758.9%estimated ± 3.4 pp, low confidence
42Qwen3.8-Flash-Next58.7%measured
43Qwen3.8-Omni-Flash57.8%measured
44Qwen3.8 Max56.6%measured
45Claude Haiku 4.556.3%estimated ± 7.2 pp, low confidence
46Composer 2.556.3%estimated ± 7.2 pp, low confidence
47Gemini 3.1 Flash-Lite56.3%estimated ± 7.2 pp, low confidence
48Gemini 3 Flash56.3%estimated ± 7.2 pp, low confidence
49GPT-5.2-Codex56.3%estimated ± 7.2 pp, low confidence
50GPT-5.3 Codex56.3%estimated ± 7.2 pp, low confidence
51Grok 4.2056.3%estimated ± 7.2 pp, low confidence
52Laguna M.156.3%estimated ± 7.2 pp, low confidence
53Laguna XS.256.3%estimated ± 7.2 pp, low confidence
54MiMo-V2.556.3%estimated ± 7.2 pp, low confidence
55Qwen3.5 Flash56.3%estimated ± 7.2 pp, low confidence
56Qwen 3.6 Max (preview)56.3%estimated ± 7.2 pp, low confidence
57Ornith-1.5-397B56.0%measured
58Gemini 3.5 Flash55.8%estimated ± 6.4 pp, medium confidence
59DeepSeek V4 Flash 073154.4%measured
60Gemini 3.1 Pro54.2%estimated ± 4.4 pp, medium confidence
61Muse Spark 1.153.3%measured
62Grok 4.553.0%measured
63GPT-5.450.0%estimated ± 4.4 pp, medium confidence
64dots3-note Preview49.3%estimated ± 9.6 pp, medium confidence
65Gemini 3.6 Flash49.0%measured
66Ornith-1.0-397B47.2%estimated ± 9.6 pp, medium confidence
67Qwen3.7 Max46.6%estimated ± 6.4 pp, low confidence
68Seed 2.1 Pro45.6%estimated ± 9.6 pp, medium confidence
69Beam44.4%measured
70Qwen3.8-27B42.2%measured
71Seed 2.1 Turbo41.3%estimated ± 9.6 pp, medium confidence
72Inkling40.8%estimated ± 4.2 pp, medium confidence
73Claude Opus 4.540.6%estimated ± 9.6 pp, medium confidence
74Laguna S 2.140.4%measured
75Kimi K2.637.0%estimated ± 6.4 pp, low confidence
76Quasar 438B35.7%estimated ± 6.4 pp, low confidence
77Apodex 1.134.7%estimated ± 6.4 pp, low confidence
78Apodex 1.1 Mini34.7%estimated ± 6.4 pp, low confidence
79Kimi K2.7 Code34.7%estimated ± 6.4 pp, low confidence
80MiMo-V2.5-Pro33.4%estimated ± 6.4 pp, low confidence
81Hy330.4%estimated ± 6.4 pp, low confidence
82Hy3 Preview30.4%estimated ± 6.4 pp, low confidence
83Ornith-1.0-35B30.3%estimated ± 9.6 pp, low confidence
84Muse Spark30.0%estimated ± 6.4 pp, low confidence
85MiniMax M329.9%estimated ± 6.4 pp, low confidence
86Ornith-1.5-9B27.8%estimated ± 9.6 pp, low confidence
87GPT-5.4 nano24.7%estimated ± 6.4 pp, low confidence
88Qwen3.7 Plus24.3%estimated ± 6.4 pp, low confidence
89GLM-5.124.1%estimated ± 6.4 pp, low confidence
90Ornith-1.0-9B22.2%estimated ± 9.6 pp, low confidence
91GPT-5.4 mini22.2%estimated ± 3.4 pp, low confidence
92Ornith-1.5-35B-A3B22.0%measured
93Claude Opus 4.621.8%estimated ± 3.4 pp, low confidence
94Qwen3.6 Plus21.7%estimated ± 6.4 pp, low confidence
95Qwen3.6-27B20.2%estimated ± 6.4 pp, low confidence
96Inkling-Small18.8%estimated ± 6.4 pp, low confidence
97MiniMax M2.718.3%estimated ± 6.4 pp, low confidence
98Ling 3.0 Flash15.1%estimated ± 6.4 pp, low confidence
99Ling 3.0 Flash FP815.1%estimated ± 6.4 pp, low confidence
100MiMo-V2-Flash13.9%estimated ± 6.4 pp, low confidence
101Claude Sonnet 4.613.7%estimated ± 3.4 pp, low confidence
102GPT-5.113.3%estimated ± 6.4 pp, low confidence
103Gemini 3.5 Flash-Lite13.2%estimated ± 6.4 pp, low confidence
104Nemotron 3 Ultra13.1%estimated ± 6.4 pp, low confidence
105Muse Glimmer 30B12.7%estimated ± 6.4 pp, low confidence
106Mistral Medium 3.5 128B10.1%estimated ± 6.4 pp, low confidence
107Kimi K2.59.9%estimated ± 6.4 pp, low confidence
108Kimi K2.5 (Reasoning)9.9%estimated ± 6.4 pp, low confidence
109Qwen3.5-122B-A10B8.8%estimated ± 6.4 pp, low confidence
110GLM-4.78.3%estimated ± 6.4 pp, low confidence
111Gemma 4 31B6.6%estimated ± 6.4 pp, low confidence
112Grok 4.35.6%estimated ± 6.4 pp, low confidence
113Qwen3.6-35B-A3B5.4%estimated ± 6.4 pp, low confidence
114o14.0%estimated ± 6.4 pp, low confidence
115Step 3.7 Flash3.9%estimated ± 6.4 pp, low confidence
116Gemma 4 26B A4B3.7%estimated ± 6.4 pp, low confidence
117GPT-5 (high)3.0%estimated ± 6.4 pp, low confidence
118Nemotron 3 Super 100B2.9%estimated ± 6.4 pp, low confidence
119o1-preview1.6%estimated ± 6.4 pp, low confidence
120Gemini 2.5 Pro1.4%estimated ± 6.4 pp, low confidence
121K-Exaone1.1%estimated ± 6.4 pp, low confidence
122Gemma 4 12B0.9%estimated ± 6.4 pp, low confidence
123GPT-OSS 120B0.8%estimated ± 6.4 pp, low confidence
124Command A+0.5%estimated ± 6.4 pp, low confidence
125Nemotron 3.5 Lightning 30B A3B NVFP40.4%estimated ± 6.4 pp, low confidence
126Mistral Small 40.4%estimated ± 6.4 pp, low confidence
127Mistral Small 4 (Reasoning)0.4%estimated ± 6.4 pp, low confidence
128Trinity-Large-Preview0.3%estimated ± 6.4 pp, low confidence
129Trinity-Large-Thinking0.3%estimated ± 6.4 pp, low confidence
130Ling 2.6 Flash0.3%estimated ± 6.4 pp, low confidence
131Gemini 1.5 Pro0.2%estimated ± 6.4 pp, low confidence
132DeepSeek V30.2%estimated ± 6.4 pp, low confidence
133Granite 4.2 8B0.1%estimated ± 6.4 pp, low confidence
134GPT-4 Turbo0.1%estimated ± 6.4 pp, low confidence
135GPT-OSS 20B0.1%estimated ± 6.4 pp, low confidence
136GPT-4.1 mini0.1%estimated ± 6.4 pp, low confidence
137Mistral Large 30.1%estimated ± 6.4 pp, low confidence
138Claude 3 Opus0.1%estimated ± 6.4 pp, low confidence
139Llama 4 Maverick0.0%estimated ± 6.4 pp, low confidence
140Celeris-10.0%estimated ± 6.4 pp, low confidence
141Nemotron 3 Nano 30B0.0%estimated ± 6.4 pp, low confidence
142Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 6.4 pp, low confidence
143Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 6.4 pp, low confidence
144GPT-4o mini0.0%estimated ± 6.4 pp, low confidence
145GPT-4.1 nano0.0%estimated ± 6.4 pp, low confidence
146Gemma 3 27B0.0%estimated ± 6.4 pp, low confidence
147Gemma 4 E4B0.0%estimated ± 6.4 pp, low confidence
148Llama 4 Scout0.0%estimated ± 6.4 pp, low confidence
149Gemma 4 E2B0.0%estimated ± 6.4 pp, low confidence
150LFM2.5-2.6B0.0%estimated ± 6.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General