benchgap
Coding

FrontierCode 1.1 Extended leaderboard

As of 2026-10-07, the highest measured score on FrontierCode 1.1 Extended is 64.5% by GPT-6 Astra. 139 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra64.5%measured
2Claude Opus 563.6%measured
3Claude Opus 5.563.6%measured
4Claude Fable 5.162.5%estimated ± 1.6 pp, low confidence
5Claude Fable 562.4%estimated ± 1.6 pp, medium confidence
6Gemini 3.8 Flash62.3%estimated ± 1.6 pp, medium confidence
7Claude Mythos 561.5%estimated ± 3.0 pp, medium confidence
8Grok 4.661.3%measured
9DeepSeek V4 Pro 081361.1%estimated ± 2.4 pp, medium confidence
10Grok 4.761.0%estimated ± 3.0 pp, medium confidence
11GPT-5.6 Sol60.6%measured
12Sakana Fugu-Ultra59.9%estimated ± 3.0 pp, medium confidence
13Muse Spark 1.359.5%estimated ± 3.0 pp, medium confidence
14Grok 4.559.4%estimated ± 1.6 pp, medium confidence
15GLM-5.359.2%estimated ± 2.4 pp, medium confidence
16Claude Sonnet 5.559.1%measured
17Hy4 preview57.8%estimated ± 3.0 pp, medium confidence
18Beam57.7%estimated ± 3.0 pp, medium confidence
19Qwen3.8 Max Preview57.7%estimated ± 3.4 pp, low confidence
20Ornith-1.5-397B57.6%estimated ± 3.0 pp, medium confidence
21Claude Opus 4.7 (Adaptive)57.4%estimated ± 3.0 pp, medium confidence
22Qwen3.8-Omni-Flash57.1%estimated ± 3.0 pp, medium confidence
23Qwen3.8-Flash-Next56.8%estimated ± 3.0 pp, low confidence
24Ornith-1.0-397B56.7%estimated ± 3.0 pp, low confidence
25dots3-note Preview56.3%estimated ± 3.0 pp, low confidence
26Atria Dawn Preview55.9%estimated ± 3.0 pp, low confidence
27Ornith-1.5-35B-A3B55.9%estimated ± 3.0 pp, low confidence
28Laguna S 2.155.8%estimated ± 3.0 pp, low confidence
29GPT-5.6 Terra55.8%measured
30Sakana Fugu55.7%estimated ± 3.0 pp, low confidence
31GPT-5.455.2%estimated ± 3.0 pp, low confidence
32Qwen3.7 Plus55.2%estimated ± 3.0 pp, low confidence
33Claude Opus 4.855.2%estimated ± 1.6 pp, medium confidence
34Claude Sonnet 555.1%estimated ± 1.6 pp, medium confidence
35Kimi K355.1%estimated ± 1.6 pp, low confidence
36GPT-5.555.1%estimated ± 1.6 pp, low confidence
37Composer 2.555.1%estimated ± 1.6 pp, low confidence
38Gemini 3.5 Flash55.1%estimated ± 1.6 pp, low confidence
39Gemini 3.6 Flash55.1%estimated ± 1.6 pp, low confidence
40GLM-5.255.1%estimated ± 1.6 pp, low confidence
41Kimi K2.7 Code55.1%estimated ± 1.6 pp, low confidence
42GPT-5.6 Luna55.1%measured
43Claude Opus 4.555.0%estimated ± 3.0 pp, low confidence
44Step 3.7 Flash54.7%estimated ± 3.0 pp, low confidence
45GPT-5.254.5%estimated ± 3.0 pp, low confidence
46GLM-554.3%estimated ± 3.0 pp, low confidence
47Claude Opus 4.653.7%estimated ± 3.0 pp, low confidence
48MAI-Thinking-153.4%estimated ± 3.0 pp, low confidence
49Muse Glimmer 30B52.8%estimated ± 3.0 pp, low confidence
50GLM-5.3-Flash52.7%estimated ± 2.4 pp, low confidence
51Qwen3.5 397B52.7%estimated ± 3.0 pp, low confidence
52Kimi K2.552.6%estimated ± 3.0 pp, low confidence
53Ornith-1.0-35B52.5%estimated ± 3.0 pp, low confidence
54Qwen3.6-35B-A3B52.1%estimated ± 3.0 pp, low confidence
55Quasar 438B51.4%estimated ± 3.4 pp, low confidence
56Laguna XS 2.151.3%estimated ± 3.0 pp, low confidence
57Ornith-1.5-9B51.2%estimated ± 3.0 pp, low confidence
58Apodex 1.151.1%estimated ± 3.4 pp, low confidence
59Apodex 1.1 Mini51.1%estimated ± 3.4 pp, low confidence
60Hy349.9%estimated ± 3.4 pp, low confidence
61Hy3 Preview49.9%estimated ± 3.4 pp, low confidence
62Ornith-1.0-9B49.1%estimated ± 3.0 pp, low confidence
63LongCat-Flash-Lite-Sparse47.9%estimated ± 3.0 pp, low confidence
64DeepSeek V4 Flash 073146.5%estimated ± 2.4 pp, low confidence
65Ling 3.0 Flash FP844.5%estimated ± 3.4 pp, low confidence
66MiMo-V2-Flash43.9%estimated ± 3.4 pp, low confidence
67Granite 4.2 30B43.6%estimated ± 3.0 pp, low confidence
68GPT-5.143.6%estimated ± 3.4 pp, low confidence
69Muse Spark 1.242.3%estimated ± 2.4 pp, low confidence
70Kimi K2.5 (Reasoning)41.8%estimated ± 3.4 pp, low confidence
71LLaDA2.2-flash41.5%estimated ± 3.0 pp, low confidence
72Qwen3.8-27B41.2%estimated ± 2.4 pp, low confidence
73Qwen3.5-122B-A10B41.0%estimated ± 3.4 pp, low confidence
74Qwen3.8 Max40.4%estimated ± 2.4 pp, low confidence
75Gemma 4 31B39.4%estimated ± 3.4 pp, low confidence
76o136.6%estimated ± 3.4 pp, low confidence
77Gemma 4 26B A4B36.3%estimated ± 3.4 pp, low confidence
78GPT-5 (high)35.2%estimated ± 3.4 pp, low confidence
79Nemotron 3 Super 100B35.1%estimated ± 3.4 pp, low confidence
80Inkling-Small34.2%estimated ± 2.4 pp, low confidence
81Claude Opus 4.733.8%estimated ± 2.4 pp, low confidence
82Muse Spark 1.133.8%estimated ± 2.4 pp, low confidence
83o1-preview32.3%estimated ± 3.4 pp, low confidence
84Granite 4.2 8B31.9%estimated ± 3.0 pp, low confidence
85Gemini 3.7 Flash31.7%estimated ± 2.4 pp, low confidence
86K-Exaone30.7%estimated ± 3.4 pp, low confidence
87Gemma 4 12B29.8%estimated ± 3.4 pp, low confidence
88GPT-OSS 120B29.3%estimated ± 3.4 pp, low confidence
89Gemini 3.1 Pro28.4%estimated ± 2.4 pp, low confidence
90Command A+27.2%estimated ± 3.4 pp, low confidence
91GPT-5.3 Codex27.1%estimated ± 2.4 pp, low confidence
92MiniCPM5-2B26.4%estimated ± 3.0 pp, low confidence
93Inkling26.4%estimated ± 2.4 pp, low confidence
94Nemotron 3.5 Lightning 30B A3B NVFP426.3%estimated ± 3.4 pp, low confidence
95Mistral Small 426.2%estimated ± 3.4 pp, low confidence
96Mistral Small 4 (Reasoning)26.2%estimated ± 3.4 pp, low confidence
97Claude Sonnet 4.626.1%estimated ± 2.4 pp, low confidence
98Trinity-Large-Preview25.4%estimated ± 3.4 pp, low confidence
99Trinity-Large-Thinking25.4%estimated ± 3.4 pp, low confidence
100Ling 2.6 Flash25.0%estimated ± 3.4 pp, low confidence
101GLM-5.124.5%estimated ± 2.4 pp, low confidence
102Kimi K2.624.2%estimated ± 2.4 pp, low confidence
103Gemini 1.5 Pro23.5%estimated ± 3.4 pp, low confidence
104DeepSeek V323.0%estimated ± 3.4 pp, low confidence
105Gemini 3.5 Flash-Lite22.4%estimated ± 2.4 pp, low confidence
106Gemini 3 Flash22.4%estimated ± 2.4 pp, low confidence
107MiniMax M322.4%estimated ± 2.4 pp, low confidence
108GPT-4 Turbo21.6%estimated ± 3.4 pp, low confidence
109Muse Spark21.6%estimated ± 2.4 pp, low confidence
110MiMo-V2.5-Pro21.0%estimated ± 2.4 pp, low confidence
111GPT-OSS 20B20.9%estimated ± 3.4 pp, low confidence
112MiniMax M2.720.7%estimated ± 2.4 pp, low confidence
113GPT-4.1 mini20.5%estimated ± 3.4 pp, low confidence
114Mistral Large 320.4%estimated ± 3.4 pp, low confidence
115Qwen3.6 Plus20.2%estimated ± 2.4 pp, low confidence
116Claude 3 Opus19.9%estimated ± 3.4 pp, low confidence
117GPT-5.4 mini19.6%estimated ± 2.4 pp, low confidence
118Qwen 3.6 Max (preview)19.4%estimated ± 2.4 pp, low confidence
119GPT-5.2-Codex18.8%estimated ± 2.4 pp, low confidence
120Grok 4.2018.6%estimated ± 2.4 pp, low confidence
121Grok 4.317.5%estimated ± 2.4 pp, low confidence
122MiMo-V2.517.0%estimated ± 2.4 pp, low confidence
123Llama 4 Maverick16.8%estimated ± 3.4 pp, low confidence
124Qwen3.6-27B15.8%estimated ± 2.4 pp, low confidence
125GPT-5.4 nano15.6%estimated ± 2.4 pp, low confidence
126GLM-4.715.1%estimated ± 2.4 pp, low confidence
127Celeris-115.0%estimated ± 3.4 pp, low confidence
128Nemotron 3 Nano 30B15.0%estimated ± 3.4 pp, low confidence
129Nemotron 3 Ultra14.7%estimated ± 2.4 pp, low confidence
130Qwen3.7 Max14.5%estimated ± 2.4 pp, low confidence
131Nemotron 3 Nano Omni 30B A3B14.4%estimated ± 3.4 pp, low confidence
132Ultravox v0.6 Llama 3.3 70B12.6%estimated ± 3.4 pp, low confidence
133Claude Haiku 4.512.2%estimated ± 2.4 pp, low confidence
134GPT-4o mini12.1%estimated ± 3.4 pp, low confidence
135Mistral Medium 3.5 128B12.0%estimated ± 2.4 pp, low confidence
136GPT-4.1 nano11.8%estimated ± 3.4 pp, low confidence
137Ling 3.0 Flash10.8%estimated ± 2.4 pp, low confidence
138Gemma 3 27B10.8%estimated ± 3.4 pp, low confidence
139Qwen3.5 Flash10.1%estimated ± 2.4 pp, low confidence
140Gemma 4 E4B10.1%estimated ± 3.4 pp, low confidence
141Llama 4 Scout8.8%estimated ± 3.4 pp, low confidence
142Gemini 3.1 Flash-Lite8.8%estimated ± 2.4 pp, low confidence
143LFM2.5-2.6B8.4%estimated ± 3.4 pp, low confidence
144Gemma 4 E2B7.8%estimated ± 3.4 pp, low confidence
145Laguna M.15.4%estimated ± 2.4 pp, low confidence
146Laguna XS.24.2%estimated ± 2.4 pp, low confidence
147Gemini 2.5 Pro3.9%estimated ± 2.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General