benchgap
Coding

Vals SWE-bench leaderboard

As of 2026-10-07, the highest measured score on Vals SWE-bench is 97.0% by Claude Opus 5. 150 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.1100.0%estimated ± 4.5 pp, medium confidence
2Gemini 4 Argon100.0%estimated ± 4.5 pp, medium confidence
3Claude Mythos 597.7%estimated ± 6.8 pp, medium confidence
4DeepSeek V4.1 Flash97.3%estimated ± 5.4 pp, low confidence
5GPT-6 Astra97.1%estimated ± 1.4 pp, low confidence
6Claude Opus 597.0%measured
7Claude Opus 5.596.9%estimated ± 1.4 pp, medium confidence
8DeepSeek V4 Pro 081396.4%measured
9GPT-5.6 Sol96.2%measured
10Grok 4.695.6%measured
11Claude Sonnet 5.595.4%estimated ± 1.4 pp, medium confidence
12GLM-5.395.4%measured
13GPT-5.6 Terra95.4%measured
14Claude Fable 595.0%measured
15Kimi K393.4%measured
16Ember-193.1%estimated ± 6.8 pp, medium confidence
17GPT-5.6 Luna93.0%measured
18MiMo-V2.6-Pro93.0%estimated ± 7.6 pp, low confidence
19Ornith-1.5-397B92.6%estimated ± 5.4 pp, medium confidence
20Grok 4.792.4%estimated ± 4.9 pp, high confidence
21Hy4 preview92.1%estimated ± 5.4 pp, medium confidence
22GLM-5.3-Flash92.0%measured
23SWE-290.9%estimated ± 2.8 pp, high confidence
24Muse Spark 1.389.3%estimated ± 5.4 pp, medium confidence
25GPT-6 Sol88.8%estimated ± 7.6 pp, low confidence
26DeepSeek V4 Flash 073188.8%measured
27Claude Opus 4.888.6%measured
28Composer 288.1%estimated ± 1.4 pp, low confidence
29Claude Opus 4.7 (Adaptive)87.7%estimated ± 5.4 pp, medium confidence
30Composer 2 Fast87.5%estimated ± 1.4 pp, low confidence
31Grok 4.586.6%measured
32Muse Spark 1.286.6%measured
33Claude Haiku 5.586.6%estimated ± 2.8 pp, high confidence
34Qwen3.8 Max Preview86.4%estimated ± 5.4 pp, medium confidence
35Qwen3.8-27B86.0%measured
36Qwen3.8 Max85.6%measured
37GPT-6 Luna85.0%estimated ± 7.6 pp, low confidence
38GPT-6.1 Sol84.5%estimated ± 7.6 pp, low confidence
39Ling 3.1 Flash84.4%estimated ± 7.6 pp, low confidence
40dots3-note Preview83.8%estimated ± 5.4 pp, medium confidence
41Qwen3.8-Omni-Flash82.9%estimated ± 5.4 pp, medium confidence
42GLM-5.282.8%measured
43GPT-5.582.6%measured
44SWE-1.782.3%estimated ± 2.8 pp, high confidence
45GPT-5.482.3%estimated ± 1.4 pp, low confidence
46Ornith-1.0-397B82.2%estimated ± 5.4 pp, medium confidence
47Inkling-Small82.2%measured
48Mistral Large 482.1%estimated ± 6.3 pp, low confidence
49Qwen3.8-Flash-Next82.1%estimated ± 5.4 pp, medium confidence
50Claude Opus 4.782.0%measured
51Muse Spark 1.182.0%measured
52Claude Opus 4.681.6%estimated ± 1.4 pp, medium confidence
53Atria Dawn Preview81.4%estimated ± 7.5 pp, medium confidence
54Laguna S 2.181.2%estimated ± 7.5 pp, medium confidence
55Seed 2.1 Pro81.0%estimated ± 5.4 pp, medium confidence
56MiMo-V2.6-Flash80.9%estimated ± 7.6 pp, low confidence
57Gemini 3.7 Flash80.8%measured
58GPT-5.280.2%estimated ± 6.3 pp, medium confidence
59Claude Opus 4.6 (Adaptive)80.2%estimated ± 6.3 pp, medium confidence
60Ornith-1.5-35B-A3B80.2%estimated ± 5.4 pp, medium confidence
61Gemini 3.8 Flash80.0%measured
62Claude Sonnet 579.6%measured
63Composer 2.579.6%measured
64Gemini 3.6 Flash79.6%measured
65Gemini 3.1 Pro78.8%measured
66Gemini 3.5 Flash78.8%measured
67Quasar 438B78.5%estimated ± 5.4 pp, medium confidence
68Apodex 1.178.2%estimated ± 5.4 pp, medium confidence
69Apodex 1.1 Mini78.2%estimated ± 5.4 pp, medium confidence
70Kimi K2.7 Code78.2%measured
71GPT-5.3 Codex78.0%measured
72Sakana Fugu78.0%estimated ± 5.3 pp, low confidence
73Step 5 Preview77.8%estimated ± 5.3 pp, low confidence
74Sakana Fugu-Ultra77.7%estimated ± 5.3 pp, low confidence
75Inkling77.6%measured
76Seed 2.1 Turbo77.5%estimated ± 5.4 pp, medium confidence
77Kimi K2.577.5%estimated ± 1.4 pp, medium confidence
78Claude Sonnet 4.677.4%measured
79Claude Opus 4.577.0%estimated ± 5.4 pp, medium confidence
80Hy376.8%estimated ± 5.4 pp, medium confidence
81GLM-5.176.4%measured
82Gemma 4 31B76.2%estimated ± 1.4 pp, medium confidence
83Kimi K2.676.2%measured
84GLM-576.0%estimated ± 1.4 pp, medium confidence
85Qwen3.7 Plus75.4%estimated ± 5.3 pp, low confidence
86Gemini 3.5 Flash-Lite75.0%measured
87Gemini 3 Flash75.0%measured
88MiniMax M375.0%measured
89Beam74.6%estimated ± 5.3 pp, low confidence
90Grok 474.6%estimated ± 1.4 pp, medium confidence
91Muse Spark74.4%measured
92BTL-474.2%estimated ± 6.8 pp, medium confidence
93MiMo-V2.5-Pro74.0%measured
94GPT-OSS 120B73.9%estimated ± 1.4 pp, medium confidence
95DeepSeek V3.273.9%estimated ± 1.4 pp, medium confidence
96MiniMax M2.773.8%measured
97MiMo-V2-Pro73.6%estimated ± 6.8 pp, medium confidence
98GLM-5 (Reasoning)73.6%estimated ± 6.3 pp, medium confidence
99GPT-OSS 20B73.5%estimated ± 1.4 pp, low confidence
100Qwen3.6 Plus73.4%measured
101Claude Sonnet 4.5 Thinking73.3%estimated ± 6.3 pp, medium confidence
102GPT-5.1-Codex-Max73.2%estimated ± 6.3 pp, medium confidence
103GPT-5.4 mini73.0%measured
104Qwen 3.6 Max (preview)72.8%measured
105Claude Opus 4.5 Thinking72.6%estimated ± 6.3 pp, medium confidence
106Claude Sonnet 4.572.5%estimated ± 6.8 pp, medium confidence
107GPT-5.2-Codex72.4%measured
108Grok 4.2072.2%measured
109Ling 3.0 Flash VL71.9%estimated ± 7.6 pp, low confidence
110Grok 4.371.4%measured
111Qwen3.5 397B71.1%estimated ± 6.8 pp, medium confidence
112MiMo-V2.571.0%measured
113Qwen3.5 Plus70.5%estimated ± 6.3 pp, medium confidence
114MiMo-V2-Flash70.2%estimated ± 5.4 pp, medium confidence
115MiniMax M2.570.1%estimated ± 6.3 pp, medium confidence
116Qwen3.6-27B70.0%measured
117Gemini 3 Pro69.9%estimated ± 6.3 pp, medium confidence
118GPT-5.169.8%estimated ± 5.4 pp, medium confidence
119GPT-5 mini69.8%estimated ± 6.3 pp, medium confidence
120GPT-5.4 nano69.8%measured
121GLM-4.769.4%measured
122GPT-5.1-Codex69.3%estimated ± 6.3 pp, medium confidence
123Muse Glimmer 30B69.3%estimated ± 5.3 pp, low confidence
124MiMo-V2-Omni69.2%estimated ± 6.8 pp, medium confidence
125Nemotron 3 Ultra69.0%measured
126Qwen3.7 Max68.8%measured
127Claude 4.1 Opus68.8%estimated ± 6.8 pp, medium confidence
128Claude Haiku 4.5 Thinking68.5%estimated ± 6.3 pp, medium confidence
129Kimi K2.5 (Reasoning)67.9%estimated ± 5.4 pp, medium confidence
130MAI-Thinking-167.4%estimated ± 6.8 pp, medium confidence
131Qwen3.5-122B-A10B67.1%estimated ± 5.4 pp, medium confidence
132Ornith-1.0-35B66.8%estimated ± 5.4 pp, low confidence
133Claude Haiku 4.566.6%measured
134Mistral Medium 3.5 128B66.4%measured
135Claude 4 Sonnet66.3%estimated ± 6.8 pp, medium confidence
136MAI-Code-1.1-Flash66.2%estimated ± 6.8 pp, medium confidence
137Qwen3.5-27B65.9%estimated ± 6.8 pp, medium confidence
138Hy3 Preview65.9%estimated ± 5.3 pp, low confidence
139A.X K265.5%estimated ± 5.3 pp, low confidence
140DeepSeek V3 032465.4%estimated ± 7.6 pp, low confidence
141Ling 3.0 Flash65.2%measured
142North Mini Code65.1%estimated ± 7.6 pp, low confidence
143DeepSeek V3.2 (Thinking)65.0%estimated ± 6.3 pp, medium confidence
144Ling 3.0 Flash FP864.4%estimated ± 5.3 pp, low confidence
145Qwen3.5 Flash64.4%measured
146Qwen3 Max64.1%estimated ± 6.3 pp, medium confidence
147Ornith-1.5-9B63.9%estimated ± 5.4 pp, low confidence
148Laguna XS 2.163.8%estimated ± 6.8 pp, medium confidence
149GLM-4.663.8%estimated ± 6.3 pp, medium confidence
150Grok Code Fast 163.7%estimated ± 6.8 pp, medium confidence
151Solar Pro 463.4%estimated ± 6.8 pp, medium confidence
152Solar Open 263.2%estimated ± 6.8 pp, medium confidence
153Gemini 3.1 Flash-Lite62.8%measured
154o162.7%estimated ± 5.4 pp, medium confidence
155Grok 4.1 Fast (Reasoning)62.6%estimated ± 6.3 pp, medium confidence
156Step 3.7 Flash62.6%estimated ± 5.4 pp, medium confidence
157Gemma 4 26B A4B62.4%estimated ± 5.4 pp, medium confidence
158Grok 4 Fast (Reasoning)61.8%estimated ± 6.3 pp, medium confidence
159Qwen3.5-35B-A3B61.5%estimated ± 6.8 pp, medium confidence
160GPT-5 (high)61.3%estimated ± 5.4 pp, medium confidence
161Nemotron 3 Super 100B61.2%estimated ± 5.4 pp, medium confidence
162Granite 4.2 30B61.2%estimated ± 5.3 pp, low confidence
163LongCat-Flash-Lite-Sparse60.1%estimated ± 6.8 pp, medium confidence
164Qwen3.6-35B-A3B59.7%estimated ± 5.4 pp, low confidence
165Mercury 2.559.4%estimated ± 5.3 pp, low confidence
166o1-preview58.5%estimated ± 5.4 pp, medium confidence
167K-EXAONE 2.057.9%estimated ± 5.3 pp, low confidence
168Laguna M.157.6%measured
169K-Exaone57.1%estimated ± 5.4 pp, low confidence
170Ornith-1.0-9B56.5%estimated ± 5.4 pp, low confidence
171Gemma 4 12B56.2%estimated ± 5.4 pp, low confidence
172Laguna XS.255.2%measured
173Granite 4.2 8B54.4%estimated ± 5.3 pp, low confidence
174Gemini 2.5 Pro54.4%measured
175Command A+53.9%estimated ± 5.4 pp, low confidence
176Mistral Small 453.0%estimated ± 5.4 pp, low confidence
177Mistral Small 4 (Reasoning)53.0%estimated ± 5.4 pp, low confidence
178Trinity-Large-Preview52.4%estimated ± 5.4 pp, low confidence
179Trinity-Large-Thinking52.4%estimated ± 5.4 pp, low confidence
180Gemini 1.5 Pro50.8%estimated ± 5.4 pp, low confidence
181DeepSeek V350.4%estimated ± 5.4 pp, low confidence
182Ternary Bonsai 2 27B50.0%estimated ± 6.8 pp, low confidence
183GPT-4 Turbo49.2%estimated ± 5.4 pp, low confidence
184Solar Pro 348.4%estimated ± 7.6 pp, low confidence
185GPT-4.1 mini48.3%estimated ± 5.4 pp, low confidence
186Mistral Large 348.2%estimated ± 5.4 pp, low confidence
187Claude 3 Opus47.8%estimated ± 5.4 pp, low confidence
188Ling 3.0 Tiny46.7%estimated ± 7.6 pp, low confidence
189Llama 4 Maverick45.4%estimated ± 5.4 pp, low confidence
190Celeris-144.0%estimated ± 5.4 pp, low confidence
191Nemotron 3 Nano 30B44.0%estimated ± 5.4 pp, low confidence
192Ultravox v0.6 Llama 3.3 70B42.2%estimated ± 5.4 pp, low confidence
193GPT-4o mini41.8%estimated ± 5.4 pp, low confidence
194GPT-4.1 nano41.6%estimated ± 5.4 pp, low confidence
195GPT-4.141.4%estimated ± 6.8 pp, low confidence
196Gemma 3 27B40.8%estimated ± 5.4 pp, low confidence
197Nemotron 3 Nano Omni 30B A3B40.7%estimated ± 5.3 pp, low confidence
198Gemma 4 E4B40.3%estimated ± 5.4 pp, low confidence
199ZAYA1-74B-Preview39.5%estimated ± 6.8 pp, low confidence
200Llama 4 Scout39.4%estimated ± 5.4 pp, low confidence
201LFM2.5-2.6B39.1%estimated ± 5.4 pp, low confidence
202Gemma 4 E2B38.7%estimated ± 5.4 pp, low confidence
203Nemotron 3.5 Lightning 30B A3B NVFP438.4%estimated ± 5.3 pp, low confidence
204o3-mini34.1%estimated ± 6.8 pp, low confidence
205LLaDA2.2-flash34.1%estimated ± 6.8 pp, low confidence
206Claude 3.5 Sonnet33.7%estimated ± 6.8 pp, low confidence
207Agents-A1-4B31.6%estimated ± 5.3 pp, low confidence
208Ling 2.6 Flash21.8%estimated ± 5.3 pp, low confidence
209MiniCPM5-2B19.4%estimated ± 5.3 pp, low confidence
210Granite 4.2 3B12.8%estimated ± 5.3 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General