benchgap
Coding

FrontierCode 1.1 Main leaderboard

As of 2026-10-07, the highest measured score on FrontierCode 1.1 Main is 54.4% by Claude Opus 5.5. 190 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.157.9%estimated ± 1.6 pp, low confidence
2Muse Spark 1.354.9%estimated ± 3.7 pp, low confidence
3Grok 4.654.7%estimated ± 1.6 pp, low confidence
4Claude Opus 5.554.4%measured
5Gemini 4 Argon53.8%estimated ± 2.8 pp, medium confidence
6DeepSeek V4.1 Flash53.5%estimated ± 3.7 pp, low confidence
7Claude Fable 553.5%measured
8Claude Opus 553.4%measured
9GPT-6 Astra53.3%measured
10Pareto 26.953.3%estimated ± 3.7 pp, low confidence
11Claude Mythos 553.2%estimated ± 2.7 pp, medium confidence
12Gemini 3.8 Flash52.8%estimated ± 1.6 pp, medium confidence
13Ember-151.6%estimated ± 2.7 pp, medium confidence
14GPT-6.1 Sol50.9%estimated ± 3.7 pp, low confidence
15MiMo-V2.6-Pro50.9%estimated ± 3.7 pp, low confidence
16GPT-5.6 Sol50.5%estimated ± 1.6 pp, medium confidence
17SWE-250.0%measured
18Grok 4.749.9%estimated ± 3.7 pp, low confidence
19Grok 4.549.9%estimated ± 1.6 pp, medium confidence
20Pareto 26.10 Preview48.8%estimated ± 3.7 pp, low confidence
21GPT-5.6 Terra47.9%estimated ± 1.6 pp, medium confidence
22Mistral Large 447.7%estimated ± 4.3 pp, low confidence
23GPT-6 Sol47.6%estimated ± 3.7 pp, low confidence
24Sakana Fugu-Ultra47.1%estimated ± 5.5 pp, medium confidence
25MiMo-V2.6-Flash46.6%estimated ± 3.7 pp, low confidence
26Claude Opus 4.846.5%measured
27Step 5 Preview46.4%estimated ± 3.7 pp, low confidence
28Claude Haiku 5.546.4%measured
29Claude Opus 4.7 (Adaptive)46.3%estimated ± 2.7 pp, medium confidence
30Claude Sonnet 5.546.2%measured
31GLM-5.345.6%estimated ± 3.7 pp, low confidence
32GPT-6 Luna45.3%estimated ± 3.7 pp, low confidence
33GPT-5.6 Luna43.9%estimated ± 1.6 pp, medium confidence
34Gemini 3.7 Flash43.6%measured
35Kimi K343.6%estimated ± 1.6 pp, medium confidence
36Ornith-1.5-397B43.1%estimated ± 2.7 pp, medium confidence
37Hy4 preview43.0%estimated ± 3.7 pp, low confidence
38GPT-5.543.0%measured
39Claude Sonnet 542.7%measured
40Ling 3.1 Flash42.7%estimated ± 6.6 pp, medium confidence
41Qwen3.8 Max Preview42.6%estimated ± 3.8 pp, high confidence
42SWE-1.742.3%measured
43GLM-5.3-Flash42.1%estimated ± 3.7 pp, low confidence
44Gemini 3.1 Pro40.7%estimated ± 2.8 pp, low confidence
45GPT-5.440.7%estimated ± 2.8 pp, low confidence
46GPT-5.3 Codex40.7%estimated ± 2.7 pp, medium confidence
47Atria Dawn Preview39.9%estimated ± 5.5 pp, medium confidence
48Sakana Fugu39.6%estimated ± 5.5 pp, medium confidence
49Composer 2.538.9%estimated ± 1.6 pp, low confidence
50Claude Opus 4.738.5%measured
51Muse Spark 1.238.2%estimated ± 3.7 pp, low confidence
52GLM-5.237.9%estimated ± 1.6 pp, low confidence
53Qwen3.8-Flash-Next37.7%estimated ± 3.7 pp, low confidence
54Qwen3.8-Omni-Flash36.9%estimated ± 3.7 pp, low confidence
55Gemini 3.6 Flash36.5%estimated ± 1.6 pp, low confidence
56Qwen3.8 Max35.8%estimated ± 3.7 pp, low confidence
57Ornith-1.0-397B33.1%estimated ± 2.7 pp, medium confidence
58Kimi K2.7 Code33.0%estimated ± 1.6 pp, low confidence
59Muse Spark 1.133.0%estimated ± 3.7 pp, low confidence
60Gemini 3.5 Flash32.2%estimated ± 1.6 pp, low confidence
61Quasar 438B30.4%estimated ± 3.8 pp, high confidence
62Apodex 1.1 Mini30.0%estimated ± 3.8 pp, high confidence
63MiMo-V2.5-Pro29.4%estimated ± 3.8 pp, high confidence
64Hy328.2%estimated ± 3.8 pp, high confidence
65Claude Opus 4.528.1%estimated ± 2.7 pp, medium confidence
66Beam28.1%estimated ± 2.7 pp, medium confidence
67Qwen 3.6 Max (preview)27.6%estimated ± 5.2 pp, low confidence
68DeepSeek V4 Pro 081327.1%estimated ± 2.7 pp, medium confidence
69Claude Opus 4.6 (Adaptive)27.0%estimated ± 4.3 pp, medium confidence
70GPT-5.4 mini27.0%measured
71Claude Opus 4.626.9%measured
72MiniMax M326.8%estimated ± 2.7 pp, medium confidence
73Qwen3.7 Max26.5%estimated ± 2.7 pp, medium confidence
74GPT-5.4 nano25.8%estimated ± 3.8 pp, medium confidence
75Inkling-Small25.8%estimated ± 2.7 pp, medium confidence
76Kimi K2.625.8%estimated ± 2.7 pp, medium confidence
77GLM-5.125.6%estimated ± 3.8 pp, medium confidence
78MiMo-V2.525.1%estimated ± 5.2 pp, low confidence
79GPT-5.225.1%estimated ± 2.7 pp, medium confidence
80Claude Sonnet 4.624.3%measured
81Qwen3.8-27B24.3%estimated ± 3.7 pp, low confidence
82MiniMax M2.723.1%estimated ± 3.8 pp, medium confidence
83Laguna S 2.123.0%estimated ± 3.7 pp, low confidence
84DeepSeek V4 Flash 073121.8%estimated ± 2.7 pp, low confidence
85Ornith-1.5-35B-A3B21.8%estimated ± 2.7 pp, low confidence
86Ling 3.0 Flash21.6%estimated ± 3.8 pp, medium confidence
87Ling 3.0 Flash FP821.6%estimated ± 3.8 pp, medium confidence
88Qwen3.6 Plus21.1%estimated ± 2.7 pp, low confidence
89GPT-5.120.8%estimated ± 3.8 pp, medium confidence
90Gemini 3.5 Flash-Lite20.7%estimated ± 3.8 pp, medium confidence
91BTL-419.9%estimated ± 2.7 pp, low confidence
92dots3-note Preview19.9%estimated ± 2.7 pp, low confidence
93MiMo-V2-Pro18.6%estimated ± 2.7 pp, low confidence
94GLM-518.0%estimated ± 2.7 pp, low confidence
95Apodex 1.117.7%estimated ± 2.7 pp, low confidence
96Qwen3.7 Plus17.7%estimated ± 2.7 pp, low confidence
97Inkling17.4%estimated ± 2.7 pp, low confidence
98Mistral Medium 3.5 128B17.4%estimated ± 2.7 pp, low confidence
99GPT-5.2-Codex17.3%estimated ± 4.3 pp, low confidence
100Gemma 4 31B16.9%estimated ± 3.8 pp, medium confidence
101Muse Spark16.8%estimated ± 2.7 pp, low confidence
102Qwen3.5 Flash16.6%estimated ± 5.2 pp, low confidence
103Claude Sonnet 4.516.3%estimated ± 2.7 pp, low confidence
104Qwen3.6-27B16.3%estimated ± 2.7 pp, low confidence
105Grok 4.316.2%estimated ± 3.8 pp, medium confidence
106Kimi K2.515.2%estimated ± 2.7 pp, low confidence
107Kimi K2.5 (Reasoning)15.2%estimated ± 2.7 pp, low confidence
108Grok 4.2014.9%estimated ± 2.7 pp, low confidence
109o114.8%estimated ± 3.8 pp, medium confidence
110Step 3.7 Flash14.7%estimated ± 3.8 pp, medium confidence
111Gemma 4 26B A4B14.6%estimated ± 3.8 pp, medium confidence
112GPT-5 (high)13.8%estimated ± 3.8 pp, medium confidence
113Nemotron 3 Super 100B13.8%estimated ± 3.8 pp, medium confidence
114Qwen3.5 397B13.6%estimated ± 2.7 pp, low confidence
115GLM-4.513.3%estimated ± 6.2 pp, low confidence
116Muse Glimmer 30B13.1%estimated ± 2.7 pp, low confidence
117Ornith-1.0-35B12.1%estimated ± 2.7 pp, low confidence
118o1-preview11.9%estimated ± 3.8 pp, medium confidence
119K-Exaone11.0%estimated ± 3.8 pp, medium confidence
120Gemma 4 12B10.5%estimated ± 3.8 pp, medium confidence
121MiMo-V2-Omni10.4%estimated ± 2.7 pp, low confidence
122GPT-OSS 120B10.3%estimated ± 3.8 pp, medium confidence
123Laguna M.19.9%estimated ± 2.7 pp, low confidence
124GLM-5 (Reasoning)9.8%estimated ± 4.3 pp, low confidence
125Claude 4.1 Opus9.7%estimated ± 2.7 pp, low confidence
126Hy3 Preview9.5%estimated ± 2.7 pp, low confidence
127Claude Sonnet 4.5 Thinking9.4%estimated ± 4.3 pp, low confidence
128GPT-5.1-Codex-Max9.2%estimated ± 4.3 pp, low confidence
129Command A+9.1%estimated ± 3.8 pp, medium confidence
130Mistral Small 48.6%estimated ± 3.8 pp, medium confidence
131Mistral Small 4 (Reasoning)8.6%estimated ± 3.8 pp, medium confidence
132Claude Opus 4.5 Thinking8.5%estimated ± 4.3 pp, low confidence
133GLM-4.78.4%estimated ± 2.7 pp, low confidence
134Gemini 3 Flash8.3%estimated ± 4.3 pp, low confidence
135Trinity-Large-Preview8.3%estimated ± 3.8 pp, medium confidence
136Trinity-Large-Thinking8.3%estimated ± 3.8 pp, medium confidence
137Ling 2.6 Flash8.1%estimated ± 3.8 pp, medium confidence
138MAI-Thinking-17.9%estimated ± 2.7 pp, low confidence
139MiMo-V2-Flash7.7%estimated ± 2.7 pp, low confidence
140Qwen3.6-35B-A3B7.7%estimated ± 2.7 pp, low confidence
141Claude Haiku 4.57.6%estimated ± 2.7 pp, low confidence
142Gemini 1.5 Pro7.4%estimated ± 3.8 pp, medium confidence
143Claude 4 Sonnet6.7%estimated ± 2.7 pp, low confidence
144GPT-4 Turbo6.6%estimated ± 3.8 pp, medium confidence
145MAI-Code-1.1-Flash6.5%estimated ± 2.7 pp, low confidence
146GPT-OSS 20B6.3%estimated ± 3.8 pp, medium confidence
147Qwen3.5 Plus6.3%estimated ± 4.3 pp, low confidence
148Qwen3.5-27B6.2%estimated ± 2.7 pp, low confidence
149Mistral Large 36.1%estimated ± 3.8 pp, medium confidence
150MiniMax M2.55.9%estimated ± 4.3 pp, low confidence
151Claude 3 Opus5.9%estimated ± 3.8 pp, medium confidence
152Qwen3.5-122B-A10B5.7%estimated ± 2.7 pp, low confidence
153Gemini 3 Pro5.7%estimated ± 4.3 pp, low confidence
154GPT-5 mini5.6%estimated ± 4.3 pp, low confidence
155Nemotron 3 Ultra5.6%estimated ± 2.7 pp, low confidence
156GPT-5.1-Codex5.2%estimated ± 4.3 pp, low confidence
157Llama 4 Maverick4.8%estimated ± 3.8 pp, medium confidence
158Claude Haiku 4.5 Thinking4.5%estimated ± 4.3 pp, low confidence
159Laguna XS 2.14.5%estimated ± 2.7 pp, low confidence
160Grok Code Fast 14.4%estimated ± 2.7 pp, low confidence
161Ornith-1.5-9B4.2%estimated ± 2.7 pp, low confidence
162Solar Pro 44.2%estimated ± 2.7 pp, low confidence
163Celeris-14.2%estimated ± 3.8 pp, medium confidence
164Nemotron 3 Nano 30B4.1%estimated ± 3.8 pp, medium confidence
165Solar Open 24.0%estimated ± 2.7 pp, low confidence
166Nemotron 3 Nano Omni 30B A3B3.9%estimated ± 3.8 pp, medium confidence
167Laguna XS.23.5%estimated ± 2.7 pp, low confidence
168Ultravox v0.6 Llama 3.3 70B3.4%estimated ± 3.8 pp, medium confidence
169GPT-4o mini3.2%estimated ± 3.8 pp, medium confidence
170Ornith-1.0-9B3.2%estimated ± 2.7 pp, low confidence
171GPT-4.1 nano3.1%estimated ± 3.8 pp, medium confidence
172Qwen3.5-35B-A3B3.0%estimated ± 2.7 pp, low confidence
173Gemma 3 27B2.8%estimated ± 3.8 pp, medium confidence
174Gemma 4 E4B2.6%estimated ± 3.8 pp, medium confidence
175K-EXAONE 2.02.4%estimated ± 2.7 pp, low confidence
176LongCat-Flash-Lite-Sparse2.4%estimated ± 2.7 pp, low confidence
177Llama 4 Scout2.2%estimated ± 3.8 pp, medium confidence
178LFM2.5-2.6B2.1%estimated ± 3.8 pp, medium confidence
179Gemma 4 E2B2.0%estimated ± 3.8 pp, medium confidence
180DeepSeek V3.2 (Thinking)1.9%estimated ± 4.3 pp, low confidence
181Qwen3 Max1.3%estimated ± 4.3 pp, low confidence
182GLM-4.61.2%estimated ± 4.3 pp, low confidence
183Gemini 2.5 Pro0.8%estimated ± 2.7 pp, low confidence
184Grok 4.1 Fast (Reasoning)0.4%estimated ± 4.3 pp, low confidence
185Ternary Bonsai 2 27B0.4%estimated ± 2.7 pp, low confidence
186Granite 4.2 30B0.2%estimated ± 2.7 pp, low confidence
187Ling 3.0 Flash VL0.1%estimated ± 6.6 pp, low confidence
188GPT-4.10.1%estimated ± 2.7 pp, low confidence
189ZAYA1-74B-Preview0.1%estimated ± 2.7 pp, low confidence
190Nemotron 3.5 Lightning 30B A3B NVFP40.1%estimated ± 2.7 pp, low confidence
191o3-mini0.0%estimated ± 2.7 pp, low confidence
192LLaDA2.2-flash0.0%estimated ± 2.7 pp, low confidence
193Claude 3.5 Sonnet0.0%estimated ± 2.7 pp, low confidence
194Granite 4.2 8B0.0%estimated ± 2.7 pp, low confidence
195A.X K20.0%estimated ± 6.6 pp, low confidence
196MiniCPM5-2B0.0%estimated ± 2.7 pp, low confidence
197DeepSeek V30.0%estimated ± 2.7 pp, low confidence
198DeepSeek V3 03240.0%estimated ± 6.6 pp, low confidence
199North Mini Code0.0%estimated ± 6.6 pp, low confidence
200Mercury 2.50.0%estimated ± 6.6 pp, low confidence
201Gemini 3.1 Flash-Lite0.0%estimated ± 4.3 pp, low confidence
202GPT-4.1 mini0.0%estimated ± 2.7 pp, low confidence
203Granite 4.2 3B0.0%estimated ± 6.6 pp, low confidence
204Grok 4 Fast (Reasoning)0.0%estimated ± 4.3 pp, low confidence
205Ling 3.0 Tiny0.0%estimated ± 6.6 pp, low confidence
206Solar Pro 30.0%estimated ± 6.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General