benchgap
Coding

SWE-Rebench leaderboard

As of 2026-10-07, the highest measured score on SWE-Rebench is 65.3% by Claude Opus 4.6. 158 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 4.665.3%measured
2GLM-562.8%measured
3GLM-5.162.7%measured
4Gemini 4 Argon62.3%estimated ± 6.5 pp, low confidence
5Mistral Large 462.2%estimated ± 6.5 pp, low confidence
6Claude Opus 4.762.2%estimated ± 6.5 pp, low confidence
7GPT-5.562.2%estimated ± 6.5 pp, low confidence
8GPT-5.462.2%estimated ± 6.5 pp, low confidence
9Claude Opus 4.6 (Adaptive)62.1%estimated ± 6.5 pp, low confidence
10Gemini 3.5 Flash62.0%estimated ± 6.5 pp, low confidence
11GPT-5.4 mini62.0%estimated ± 6.5 pp, low confidence
12Claude Fable 561.8%estimated ± 3.6 pp, low confidence
13Claude Mythos 561.8%estimated ± 3.6 pp, low confidence
14Claude Opus 4.7 (Adaptive)61.8%estimated ± 3.6 pp, low confidence
15Claude Opus 4.861.8%estimated ± 3.6 pp, low confidence
16Claude Opus 561.8%estimated ± 3.6 pp, low confidence
17Claude Sonnet 561.8%estimated ± 3.6 pp, low confidence
18Ember-161.8%estimated ± 3.6 pp, low confidence
19Ornith-1.5-397B61.8%estimated ± 3.6 pp, low confidence
20Ornith-1.0-397B61.8%estimated ± 3.6 pp, medium confidence
21Claude Opus 4.561.8%estimated ± 3.6 pp, medium confidence
22Beam61.8%estimated ± 3.6 pp, medium confidence
23DeepSeek V4 Pro 081361.8%estimated ± 3.6 pp, medium confidence
24MiniMax M361.8%estimated ± 3.6 pp, medium confidence
25Qwen3.7 Max61.8%estimated ± 3.6 pp, medium confidence
26Inkling-Small61.8%estimated ± 3.6 pp, medium confidence
27Kimi K2.661.8%estimated ± 3.6 pp, medium confidence
28GPT-5.261.8%estimated ± 3.6 pp, medium confidence
29DeepSeek V4 Flash 073161.7%estimated ± 3.6 pp, medium confidence
30Ornith-1.5-35B-A3B61.7%estimated ± 3.6 pp, medium confidence
31Qwen3.6 Plus61.7%estimated ± 3.6 pp, medium confidence
32BTL-461.6%estimated ± 3.6 pp, medium confidence
33dots3-note Preview61.6%estimated ± 3.6 pp, medium confidence
34MiMo-V2-Pro61.3%estimated ± 3.6 pp, medium confidence
35Apodex 1.161.0%estimated ± 3.6 pp, medium confidence
36Qwen3.7 Plus61.0%estimated ± 3.6 pp, medium confidence
37GPT-5.2-Codex61.0%estimated ± 6.5 pp, low confidence
38DeepSeek V3.260.9%measured
39Inkling60.9%estimated ± 3.6 pp, medium confidence
40Mistral Medium 3.5 128B60.9%estimated ± 3.6 pp, medium confidence
41Claude Sonnet 4.660.7%measured
42Muse Spark60.6%estimated ± 3.6 pp, medium confidence
43Claude Sonnet 4.560.2%estimated ± 3.6 pp, medium confidence
44Qwen3.6-27B60.2%estimated ± 3.6 pp, medium confidence
45Kimi K2.5 (Reasoning)59.3%estimated ± 3.6 pp, medium confidence
46Grok 4.2059.0%estimated ± 3.6 pp, medium confidence
47Gemini 3.1 Pro58.9%estimated ± 6.5 pp, low confidence
48Qwen3.5-27B58.9%measured
49GLM-4.758.7%measured
50Kimi K2.558.5%measured
51GPT-5.3 Codex58.2%measured
52Apodex 1.1 Mini58.1%estimated ± 8.9 pp, low confidence
53Claude Fable 5.158.1%estimated ± 8.9 pp, low confidence
54Gemini 3.6 Flash58.1%estimated ± 8.9 pp, low confidence
55Gemini 3.7 Flash58.1%estimated ± 8.9 pp, low confidence
56Gemini 3.8 Flash58.1%estimated ± 8.9 pp, low confidence
57GLM-5.258.1%estimated ± 8.9 pp, low confidence
58GLM-5.358.1%estimated ± 8.9 pp, low confidence
59GPT-5.6 Luna58.1%estimated ± 8.9 pp, low confidence
60GPT-5.6 Sol58.1%estimated ± 8.9 pp, low confidence
61GPT-5.6 Terra58.1%estimated ± 8.9 pp, low confidence
62GPT-6 Astra58.1%estimated ± 8.9 pp, low confidence
63Grok 4.558.1%estimated ± 8.9 pp, low confidence
64Grok 4.658.1%estimated ± 8.9 pp, low confidence
65Hy358.1%estimated ± 8.9 pp, low confidence
66Kimi K2.7 Code58.1%estimated ± 8.9 pp, low confidence
67Kimi K358.1%estimated ± 8.9 pp, low confidence
68Ling 3.0 Flash58.1%estimated ± 8.9 pp, low confidence
69Ling 3.0 Flash FP858.1%estimated ± 8.9 pp, low confidence
70MiMo-V2.5-Pro58.1%estimated ± 8.9 pp, low confidence
71Muse Spark 1.158.1%estimated ± 8.9 pp, low confidence
72Muse Spark 1.258.1%estimated ± 8.9 pp, low confidence
73Muse Spark 1.358.1%estimated ± 8.9 pp, low confidence
74Quasar 438B58.1%estimated ± 8.9 pp, low confidence
75Qwen3.8-27B58.1%estimated ± 8.9 pp, low confidence
76Qwen3.8-Flash-Next58.1%estimated ± 8.9 pp, low confidence
77Qwen3.8 Max Preview58.1%estimated ± 8.9 pp, low confidence
78Gemini 3.5 Flash-Lite58.1%estimated ± 8.9 pp, low confidence
79Qwen3.5 397B58.0%estimated ± 3.6 pp, medium confidence
80Composer 258.0%measured
81Muse Glimmer 30B57.8%estimated ± 3.6 pp, medium confidence
82Ornith-1.0-35B57.4%estimated ± 3.6 pp, medium confidence
83MiMo-V2-Omni57.1%estimated ± 3.6 pp, medium confidence
84Laguna M.157.1%estimated ± 3.6 pp, medium confidence
85Claude 4.1 Opus57.1%estimated ± 3.6 pp, medium confidence
86Hy3 Preview57.1%estimated ± 3.6 pp, medium confidence
87MAI-Thinking-157.0%estimated ± 3.6 pp, medium confidence
88MiMo-V2-Flash57.0%estimated ± 3.6 pp, medium confidence
89Qwen3.6-35B-A3B57.0%estimated ± 3.6 pp, medium confidence
90Claude Haiku 4.557.0%estimated ± 3.6 pp, medium confidence
91Claude 4 Sonnet57.0%estimated ± 3.6 pp, medium confidence
92MAI-Code-1.1-Flash57.0%estimated ± 3.6 pp, medium confidence
93Nemotron 3 Ultra57.0%estimated ± 3.6 pp, medium confidence
94Qwen3.5-122B-A10B57.0%estimated ± 3.6 pp, medium confidence
95Claude 3.5 Sonnet57.0%estimated ± 3.6 pp, low confidence
96DeepSeek V357.0%estimated ± 3.6 pp, low confidence
97Gemini 2.5 Pro57.0%estimated ± 3.6 pp, low confidence
98GPT-4.157.0%estimated ± 3.6 pp, low confidence
99GPT-4.1 mini57.0%estimated ± 3.6 pp, low confidence
100Granite 4.2 30B57.0%estimated ± 3.6 pp, low confidence
101Granite 4.2 8B57.0%estimated ± 3.6 pp, low confidence
102Grok Code Fast 157.0%estimated ± 3.6 pp, medium confidence
103K-EXAONE 2.057.0%estimated ± 3.6 pp, low confidence
104Laguna XS.257.0%estimated ± 3.6 pp, medium confidence
105Laguna XS 2.157.0%estimated ± 3.6 pp, medium confidence
106LLaDA2.2-flash57.0%estimated ± 3.6 pp, low confidence
107LongCat-Flash-Lite-Sparse57.0%estimated ± 3.6 pp, low confidence
108MiniCPM5-2B57.0%estimated ± 3.6 pp, low confidence
109Nemotron 3.5 Lightning 30B A3B NVFP457.0%estimated ± 3.6 pp, low confidence
110o3-mini57.0%estimated ± 3.6 pp, low confidence
111Ornith-1.0-9B57.0%estimated ± 3.6 pp, medium confidence
112Ornith-1.5-9B57.0%estimated ± 3.6 pp, medium confidence
113Solar Open 257.0%estimated ± 3.6 pp, medium confidence
114Solar Pro 457.0%estimated ± 3.6 pp, medium confidence
115Ternary Bonsai 2 27B57.0%estimated ± 3.6 pp, low confidence
116ZAYA1-74B-Preview57.0%estimated ± 3.6 pp, low confidence
117Qwen3.5-35B-A3B53.7%measured
118GPT-5.4 nano52.2%estimated ± 6.5 pp, low confidence
119MiniMax M2.751.9%measured
120GPT-5.148.8%estimated ± 6.5 pp, low confidence
121GLM-5 (Reasoning)45.2%estimated ± 6.5 pp, low confidence
122Claude Sonnet 4.5 Thinking42.8%estimated ± 6.5 pp, low confidence
123Gemma 4 31B41.6%measured
124GPT-5.1-Codex-Max41.1%estimated ± 6.5 pp, low confidence
125Claude Opus 4.5 Thinking34.7%estimated ± 6.5 pp, low confidence
126Gemini 3 Flash32.8%estimated ± 6.5 pp, low confidence
127GPT-5 (high)32.3%estimated ± 6.5 pp, low confidence
128Qwen3.5 Plus12.4%estimated ± 6.5 pp, low confidence
129Grok 4.311.2%estimated ± 8.9 pp, low confidence
130MiniMax M2.59.3%estimated ± 6.5 pp, low confidence
131Gemini 3 Pro7.6%estimated ± 6.5 pp, low confidence
132GPT-5 mini7.3%estimated ± 6.5 pp, low confidence
133GPT-5.1-Codex4.8%estimated ± 6.5 pp, low confidence
134Claude Haiku 4.5 Thinking2.2%estimated ± 6.5 pp, low confidence
135o10.1%estimated ± 8.9 pp, low confidence
136Step 3.7 Flash0.1%estimated ± 8.9 pp, low confidence
137Gemma 4 26B A4B0.0%estimated ± 8.9 pp, low confidence
138DeepSeek V3.2 (Thinking)0.0%estimated ± 6.5 pp, low confidence
139Qwen3 Max0.0%estimated ± 6.5 pp, low confidence
140Nemotron 3 Super 100B0.0%estimated ± 8.9 pp, low confidence
141GLM-4.60.0%estimated ± 6.5 pp, low confidence
142Celeris-10.0%estimated ± 8.9 pp, low confidence
143Claude 3 Opus0.0%estimated ± 8.9 pp, low confidence
144Command A+0.0%estimated ± 8.9 pp, low confidence
145Gemini 1.5 Pro0.0%estimated ± 8.9 pp, low confidence
146Gemini 3.1 Flash-Lite0.0%estimated ± 6.5 pp, low confidence
147Gemma 3 27B0.0%estimated ± 8.9 pp, low confidence
148Gemma 4 12B0.0%estimated ± 8.9 pp, low confidence
149Gemma 4 E2B0.0%estimated ± 8.9 pp, low confidence
150Gemma 4 E4B0.0%estimated ± 8.9 pp, low confidence
151GPT-4.1 nano0.0%estimated ± 8.9 pp, low confidence
152GPT-4 Turbo0.0%estimated ± 8.9 pp, low confidence
153GPT-4o mini0.0%estimated ± 8.9 pp, low confidence
154GPT-OSS 120B0.0%estimated ± 8.9 pp, low confidence
155GPT-OSS 20B0.0%estimated ± 8.9 pp, low confidence
156Grok 4.1 Fast (Reasoning)0.0%estimated ± 6.5 pp, low confidence
157Grok 4 Fast (Reasoning)0.0%estimated ± 6.5 pp, low confidence
158K-Exaone0.0%estimated ± 8.9 pp, low confidence
159LFM2.5-2.6B0.0%estimated ± 8.9 pp, low confidence
160Ling 2.6 Flash0.0%estimated ± 8.9 pp, low confidence
161Llama 4 Maverick0.0%estimated ± 8.9 pp, low confidence
162Llama 4 Scout0.0%estimated ± 8.9 pp, low confidence
163Mistral Large 30.0%estimated ± 8.9 pp, low confidence
164Mistral Small 40.0%estimated ± 8.9 pp, low confidence
165Mistral Small 4 (Reasoning)0.0%estimated ± 8.9 pp, low confidence
166Nemotron 3 Nano 30B0.0%estimated ± 8.9 pp, low confidence
167Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 8.9 pp, low confidence
168o1-preview0.0%estimated ± 8.9 pp, low confidence
169Trinity-Large-Preview0.0%estimated ± 8.9 pp, low confidence
170Trinity-Large-Thinking0.0%estimated ± 8.9 pp, low confidence
171Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 8.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General