benchgap
Coding

FrontierSWE v2 leaderboard

As of 2026-10-07, the highest measured score on FrontierSWE v2 is 65.5% by GPT-6 Astra. 150 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra65.5%measured
2Claude Opus 5.562.3%measured
3Claude Sonnet 5.561.9%measured
4Claude Fable 5.156.3%measured
5Gemini 4 Argon55.0%measured
6Claude Opus 552.0%measured
7Claude Mythos 551.3%estimated ± 9.0 pp, medium confidence
8Claude Fable 547.0%measured
9MiMo-V2.6-Pro45.2%estimated ± 14.6 pp, low confidence
10Sakana Fugu-Ultra39.5%estimated ± 9.0 pp, medium confidence
11Step 5 Preview39.1%estimated ± 14.6 pp, low confidence
12GPT-6 Sol35.1%estimated ± 14.6 pp, low confidence
13GPT-5.6 Sol32.2%measured
14Claude Opus 4.831.5%estimated ± 9.0 pp, medium confidence
15GLM-5.330.2%measured
16Grok 4.729.5%measured
17Muse Spark 1.327.9%estimated ± 7.6 pp, medium confidence
18GPT-5.6 Terra27.1%estimated ± 7.6 pp, medium confidence
19Claude Haiku 5.527.1%estimated ± 14.6 pp, low confidence
20Kimi K325.9%measured
21GPT-6 Luna25.8%estimated ± 14.6 pp, low confidence
22Grok 4.625.3%measured
23Hy4 preview25.3%estimated ± 9.0 pp, medium confidence
24Beam24.9%estimated ± 9.0 pp, medium confidence
25GPT-6.1 Sol24.6%estimated ± 14.6 pp, low confidence
26Mistral Large 424.6%estimated ± 14.6 pp, low confidence
27Ling 3.1 Flash24.3%estimated ± 14.6 pp, low confidence
28Ornith-1.5-397B24.2%estimated ± 9.0 pp, medium confidence
29Grok 4.523.5%estimated ± 9.0 pp, medium confidence
30Claude Opus 4.7 (Adaptive)22.8%estimated ± 9.0 pp, medium confidence
31Qwen3.8-Omni-Flash21.0%estimated ± 9.0 pp, medium confidence
32Gemini 3.7 Flash20.3%measured
33Gemini 3.8 Flash19.6%measured
34Qwen3.8-Flash-Next19.6%estimated ± 9.0 pp, medium confidence
35Ornith-1.0-397B19.1%estimated ± 9.0 pp, medium confidence
36GLM-5.218.9%estimated ± 9.0 pp, medium confidence
37Qwen3.8-27B18.2%estimated ± 9.0 pp, medium confidence
38Qwen3.8 Max Preview18.1%estimated ± 14.6 pp, low confidence
39GLM-5.3-Flash18.1%measured
40Muse Spark 1.117.8%estimated ± 9.0 pp, medium confidence
41DeepSeek V4.1 Flash17.5%estimated ± 14.6 pp, low confidence
42dots3-note Preview16.9%estimated ± 9.0 pp, medium confidence
43Qwen3.7 Max16.2%estimated ± 9.0 pp, medium confidence
44Qwen3.8 Max15.8%measured
45MiMo-V2.6-Flash15.6%estimated ± 14.6 pp, low confidence
46Atria Dawn Preview14.4%estimated ± 9.0 pp, medium confidence
47Ornith-1.5-35B-A3B14.4%estimated ± 9.0 pp, medium confidence
48GPT-5.6 Luna14.4%estimated ± 7.6 pp, low confidence
49Laguna S 2.114.1%estimated ± 9.0 pp, medium confidence
50MiniMax M313.4%estimated ± 9.0 pp, medium confidence
51Sakana Fugu13.4%estimated ± 9.0 pp, medium confidence
52GPT-5.512.6%estimated ± 9.0 pp, medium confidence
53Kimi K2.612.6%estimated ± 9.0 pp, medium confidence
54GLM-5.112.3%estimated ± 9.0 pp, medium confidence
55Muse Spark 1.212.0%measured
56Claude Opus 4.711.6%estimated ± 9.5 pp, medium confidence
57Claude Sonnet 511.1%estimated ± 7.6 pp, low confidence
58GPT-5.411.0%estimated ± 9.0 pp, medium confidence
59Qwen3.7 Plus10.9%estimated ± 9.0 pp, medium confidence
60Qwen 3.6 Max (preview)10.3%estimated ± 9.0 pp, medium confidence
61MiMo-V2.5-Pro10.2%estimated ± 9.0 pp, medium confidence
62Gemini 3.6 Flash10.1%estimated ± 9.5 pp, medium confidence
63Claude Opus 4.510.0%estimated ± 9.0 pp, medium confidence
64Gemini 3.1 Pro9.7%estimated ± 9.5 pp, medium confidence
65GPT-5.3 Codex9.4%estimated ± 9.0 pp, medium confidence
66Kimi K2.7 Code9.4%estimated ± 9.5 pp, medium confidence
67Ling 3.0 Flash9.1%estimated ± 9.0 pp, medium confidence
68Qwen3.6 Plus9.1%estimated ± 9.0 pp, medium confidence
69Claude Sonnet 4.69.0%estimated ± 9.5 pp, low confidence
70Step 3.7 Flash8.6%estimated ± 9.0 pp, medium confidence
71MiniMax M2.78.4%estimated ± 9.0 pp, medium confidence
72MiMo-V2.58.2%estimated ± 9.0 pp, medium confidence
73Gemini 3 Flash8.0%estimated ± 9.5 pp, low confidence
74Inkling-Small7.8%estimated ± 9.0 pp, medium confidence
75Hy37.3%estimated ± 14.6 pp, low confidence
76Hy3 Preview7.3%estimated ± 14.6 pp, low confidence
77GPT-5.27.3%estimated ± 9.0 pp, medium confidence
78GPT-5.4 mini7.3%estimated ± 9.5 pp, low confidence
79GPT-5.2-Codex7.1%estimated ± 9.5 pp, low confidence
80DeepSeek V4 Pro 08136.9%estimated ± 9.0 pp, medium confidence
81Grok 4.36.8%estimated ± 9.5 pp, low confidence
82Gemini 3.5 Flash6.4%estimated ± 9.0 pp, medium confidence
83GLM-56.4%estimated ± 9.0 pp, medium confidence
84GPT-5.4 nano6.3%estimated ± 9.5 pp, low confidence
85GLM-4.76.2%estimated ± 9.5 pp, low confidence
86Nemotron 3 Ultra6.1%estimated ± 9.5 pp, low confidence
87Quasar 438B5.8%estimated ± 14.6 pp, low confidence
88Claude Haiku 4.55.5%estimated ± 9.5 pp, low confidence
89Mistral Medium 3.5 128B5.5%estimated ± 9.5 pp, low confidence
90Qwen3.5 Flash5.0%estimated ± 9.5 pp, low confidence
91Gemini 3.5 Flash-Lite4.8%estimated ± 9.0 pp, low confidence
92Gemini 3.1 Flash-Lite4.7%estimated ± 9.5 pp, low confidence
93Inkling4.1%measured
94MiMo-V2-Flash3.7%estimated ± 14.8 pp, low confidence
95Qwen3.6-27B3.6%estimated ± 9.0 pp, low confidence
96Composer 2.53.5%estimated ± 7.6 pp, low confidence
97GPT-5.13.5%estimated ± 14.8 pp, low confidence
98Gemini 2.5 Pro3.4%estimated ± 9.5 pp, low confidence
99Claude Opus 4.63.4%estimated ± 9.0 pp, low confidence
100Kimi K2.5 (Reasoning)2.6%estimated ± 14.8 pp, low confidence
101MAI-Thinking-12.3%estimated ± 9.0 pp, low confidence
102DeepSeek V4 Flash 07312.0%estimated ± 9.0 pp, low confidence
103Muse Spark1.6%estimated ± 9.0 pp, low confidence
104o11.0%estimated ± 14.8 pp, low confidence
105GPT-5 (high)0.7%estimated ± 14.8 pp, low confidence
106Grok 4.200.5%estimated ± 9.0 pp, low confidence
107o1-preview0.4%estimated ± 14.8 pp, low confidence
108K-Exaone0.3%estimated ± 14.8 pp, low confidence
109Gemma 4 12B0.2%estimated ± 14.8 pp, low confidence
110Ling 2.6 Flash0.1%estimated ± 14.8 pp, low confidence
111Gemini 1.5 Pro0.0%estimated ± 14.8 pp, low confidence
112GPT-4 Turbo0.0%estimated ± 14.8 pp, low confidence
113GPT-4.1 mini0.0%estimated ± 14.8 pp, low confidence
114Claude 3 Opus0.0%estimated ± 14.8 pp, low confidence
115Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 14.8 pp, low confidence
116Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 14.8 pp, low confidence
117GPT-4.1 nano0.0%estimated ± 14.8 pp, low confidence
118GPT-4o mini0.0%estimated ± 14.8 pp, low confidence
119A.X K20.0%estimated ± 14.6 pp, low confidence
120Apodex 1.10.0%estimated ± 14.6 pp, low confidence
121Apodex 1.1 Mini0.0%estimated ± 14.6 pp, low confidence
122Celeris-10.0%estimated ± 14.6 pp, low confidence
123Command A+0.0%estimated ± 14.6 pp, low confidence
124DeepSeek V30.0%estimated ± 14.6 pp, low confidence
125DeepSeek V3 03240.0%estimated ± 14.6 pp, low confidence
126Gemma 3 27B0.0%estimated ± 14.6 pp, low confidence
127Gemma 4 26B A4B0.0%estimated ± 14.6 pp, low confidence
128Gemma 4 31B0.0%estimated ± 14.6 pp, low confidence
129Gemma 4 E2B0.0%estimated ± 14.8 pp, low confidence
130Gemma 4 E4B0.0%estimated ± 14.6 pp, low confidence
131GPT-OSS 120B0.0%estimated ± 14.6 pp, low confidence
132GPT-OSS 20B0.0%estimated ± 14.6 pp, low confidence
133Granite 4.2 30B0.0%estimated ± 9.0 pp, low confidence
134Granite 4.2 3B0.0%estimated ± 14.6 pp, low confidence
135Granite 4.2 8B0.0%estimated ± 9.0 pp, low confidence
136K-EXAONE 2.00.0%estimated ± 14.6 pp, low confidence
137Kimi K2.50.0%estimated ± 9.0 pp, low confidence
138Laguna M.10.0%estimated ± 9.0 pp, low confidence
139Laguna XS.20.0%estimated ± 9.0 pp, low confidence
140Laguna XS 2.10.0%estimated ± 9.0 pp, low confidence
141LFM2.5-2.6B0.0%estimated ± 14.6 pp, low confidence
142Ling 3.0 Flash FP80.0%estimated ± 14.6 pp, low confidence
143Ling 3.0 Flash VL0.0%estimated ± 14.6 pp, low confidence
144Ling 3.0 Tiny0.0%estimated ± 14.6 pp, low confidence
145LLaDA2.2-flash0.0%estimated ± 9.0 pp, low confidence
146Llama 4 Maverick0.0%estimated ± 14.6 pp, low confidence
147Llama 4 Scout0.0%estimated ± 14.6 pp, low confidence
148LongCat-Flash-Lite-Sparse0.0%estimated ± 9.0 pp, low confidence
149Mercury 2.50.0%estimated ± 14.6 pp, low confidence
150MiniCPM5-2B0.0%estimated ± 9.0 pp, low confidence
151Mistral Large 30.0%estimated ± 14.6 pp, low confidence
152Mistral Small 40.0%estimated ± 14.6 pp, low confidence
153Mistral Small 4 (Reasoning)0.0%estimated ± 14.6 pp, low confidence
154Muse Glimmer 30B0.0%estimated ± 9.0 pp, low confidence
155Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 14.6 pp, low confidence
156Nemotron 3 Nano 30B0.0%estimated ± 14.6 pp, low confidence
157Nemotron 3 Super 100B0.0%estimated ± 14.6 pp, low confidence
158North Mini Code0.0%estimated ± 14.6 pp, low confidence
159Ornith-1.0-35B0.0%estimated ± 9.0 pp, low confidence
160Ornith-1.0-9B0.0%estimated ± 9.0 pp, low confidence
161Ornith-1.5-9B0.0%estimated ± 9.0 pp, low confidence
162Qwen3.5-122B-A10B0.0%estimated ± 14.6 pp, low confidence
163Qwen3.5 397B0.0%estimated ± 9.0 pp, low confidence
164Qwen3.6-35B-A3B0.0%estimated ± 9.0 pp, low confidence
165Solar Pro 30.0%estimated ± 14.6 pp, low confidence
166Solar Pro 40.0%estimated ± 14.6 pp, low confidence
167Trinity-Large-Preview0.0%estimated ± 14.6 pp, low confidence
168Trinity-Large-Thinking0.0%estimated ± 14.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General