benchgap
Coding

SciCode leaderboard

As of 2026-10-07, the highest measured score on SciCode is 60.1% by Sakana Fugu. 168 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.566.1%estimated ± 3.2 pp, medium confidence
2Gemini 4 Argon61.1%estimated ± 3.2 pp, medium confidence
3Claude Sonnet 5.560.3%estimated ± 3.2 pp, medium confidence
4MiMo-V2.6-Pro60.3%estimated ± 3.2 pp, medium confidence
5Sakana Fugu60.1%measured
6Step 5 Preview58.9%measured
7Sakana Fugu-Ultra58.7%measured
8Muse Spark 1.358.2%estimated ± 3.2 pp, high confidence
9Claude Fable 5.157.8%estimated ± 2.9 pp, low confidence
10GPT-6 Sol57.1%estimated ± 3.2 pp, high confidence
11Grok 4.756.9%estimated ± 3.2 pp, high confidence
12Muse Spark 1.256.9%estimated ± 3.2 pp, high confidence
13Claude Fable 556.4%estimated ± 2.9 pp, low confidence
14GPT-6 Astra56.0%estimated ± 3.2 pp, high confidence
15Gemini 3.8 Flash55.8%estimated ± 2.9 pp, low confidence
16Claude Mythos 555.3%estimated ± 3.8 pp, medium confidence
17Claude Opus 554.8%estimated ± 2.9 pp, low confidence
18Claude Haiku 5.554.5%estimated ± 3.2 pp, high confidence
19Gemini 3.7 Flash54.2%estimated ± 2.9 pp, low confidence
20GPT-6 Luna54.2%estimated ± 3.2 pp, high confidence
21Hy4 preview54.1%estimated ± 4.8 pp, high confidence
22GPT-5.454.0%estimated ± 4.7 pp, medium confidence
23Ember-153.9%estimated ± 3.8 pp, medium confidence
24GPT-6.1 Sol53.8%estimated ± 3.2 pp, high confidence
25Mistral Large 453.8%estimated ± 3.2 pp, high confidence
26Gemini 3.1 Pro53.8%estimated ± 2.9 pp, low confidence
27Ling 3.1 Flash53.7%estimated ± 3.2 pp, high confidence
28Qwen3.7 Max53.5%measured
29GPT-5.6 Luna53.2%estimated ± 3.2 pp, high confidence
30Grok 4.653.1%estimated ± 2.9 pp, low confidence
31Gemini 3.5 Flash53.1%measured
32Qwen3.8-Omni-Flash53.0%estimated ± 4.8 pp, high confidence
33Gemini 3.6 Flash52.9%estimated ± 2.9 pp, low confidence
34GPT-5.2-Codex52.7%estimated ± 2.9 pp, low confidence
35Qwen3.8 Max52.5%estimated ± 2.9 pp, low confidence
36Claude Opus 4.852.3%estimated ± 2.9 pp, low confidence
37Kimi K2.652.2%measured
38Claude Opus 4.7 (Adaptive)51.9%estimated ± 3.8 pp, medium confidence
39DeepSeek V4 Pro 081351.7%estimated ± 2.9 pp, medium confidence
40Qwen3.8 Max Preview51.7%estimated ± 3.2 pp, high confidence
41DeepSeek V4.1 Flash51.5%estimated ± 3.2 pp, high confidence
42Grok 4.551.5%estimated ± 2.9 pp, medium confidence
43DeepSeek V4 Flash 073151.3%estimated ± 2.9 pp, medium confidence
44GPT-5.3 Codex51.3%estimated ± 2.9 pp, medium confidence
45Qwen3.7 Plus51.3%measured
46Atria Dawn Preview51.2%estimated ± 4.8 pp, high confidence
47Ornith-1.5-397B51.2%estimated ± 3.8 pp, medium confidence
48Laguna S 2.151.1%estimated ± 4.8 pp, high confidence
49Composer 2.551.1%estimated ± 4.5 pp, medium confidence
50Kimi K351.1%estimated ± 2.9 pp, medium confidence
51MiMo-V2.6-Flash51.0%estimated ± 3.2 pp, high confidence
52Qwen3.8-Flash-Next50.3%estimated ± 3.2 pp, high confidence
53BTL-349.8%estimated ± 5.4 pp, medium confidence
54Beam49.7%measured
55Ornith-1.0-397B49.6%estimated ± 3.8 pp, medium confidence
56Claude Opus 4.548.9%estimated ± 3.8 pp, high confidence
57Claude Opus 4.648.9%estimated ± 3.8 pp, high confidence
58Qwen3.6 Plus48.7%estimated ± 2.9 pp, medium confidence
59Inkling-Small48.7%measured
60Kimi K2.548.7%measured
61GPT-5.248.5%estimated ± 3.8 pp, high confidence
62GPT-5.6 Terra48.5%estimated ± 2.9 pp, medium confidence
63Muse Spark 1.148.5%estimated ± 2.9 pp, medium confidence
64Hy348.3%estimated ± 3.2 pp, high confidence
65Ornith-1.5-35B-A3B48.1%estimated ± 3.8 pp, high confidence
66Gemini 3 Flash47.9%estimated ± 2.9 pp, medium confidence
67GPT-5.1-Codex47.9%estimated ± 2.9 pp, medium confidence
68Quasar 438B47.9%estimated ± 3.2 pp, high confidence
69BTL-447.8%estimated ± 3.8 pp, high confidence
70dots3-note Preview47.8%estimated ± 3.8 pp, high confidence
71Inkling47.7%estimated ± 2.9 pp, medium confidence
72MiMo-V2-Pro47.6%estimated ± 3.8 pp, high confidence
73GLM-547.5%estimated ± 3.8 pp, high confidence
74Muse Spark47.3%estimated ± 3.8 pp, high confidence
75Grok 4.347.3%measured
76GPT-5.547.3%estimated ± 2.9 pp, medium confidence
77Claude Sonnet 4.547.2%estimated ± 3.8 pp, high confidence
78Kimi K2.5 (Reasoning)47.0%estimated ± 3.8 pp, high confidence
79Qwen 3.6 Max (preview)47.0%measured
80Claude Opus 4.746.9%estimated ± 2.9 pp, medium confidence
81Qwen3.5 397B46.7%estimated ± 3.8 pp, high confidence
82Ornith-1.0-35B46.5%estimated ± 3.8 pp, high confidence
83Gemini 2.5 Pro46.1%estimated ± 3.2 pp, high confidence
84MiMo-V2-Omni46.1%estimated ± 3.8 pp, high confidence
85Claude 4.1 Opus45.9%estimated ± 3.8 pp, high confidence
86MAI-Thinking-145.5%estimated ± 3.8 pp, high confidence
87MiMo-V2-Flash45.4%estimated ± 3.8 pp, high confidence
88Apodex 1.145.3%estimated ± 3.2 pp, high confidence
89Apodex 1.1 Mini45.3%estimated ± 3.2 pp, high confidence
90Gemma 4 31B45.3%estimated ± 3.2 pp, high confidence
91Grok 4.2045.3%estimated ± 2.9 pp, medium confidence
92Claude 4 Sonnet45.1%estimated ± 3.8 pp, high confidence
93MAI-Code-1.1-Flash45.0%estimated ± 3.8 pp, high confidence
94Qwen3.5-27B44.9%estimated ± 3.8 pp, high confidence
95GPT-5.4 nano44.7%estimated ± 2.9 pp, medium confidence
96Qwen3.8-27B44.7%estimated ± 2.9 pp, medium confidence
97Nemotron 3 Ultra44.6%measured
98GPT-5.144.6%estimated ± 4.7 pp, high confidence
99Solar Pro 444.5%estimated ± 3.2 pp, high confidence
100Laguna XS 2.144.2%estimated ± 3.8 pp, high confidence
101Ling 3.0 Flash VL44.1%estimated ± 3.2 pp, high confidence
102Ornith-1.5-9B44.1%estimated ± 3.8 pp, high confidence
103Solar Open 244.0%estimated ± 3.8 pp, high confidence
104GPT-5.1-Codex-Max43.9%estimated ± 2.9 pp, low confidence
105Step 3.7 Flash43.8%estimated ± 3.2 pp, high confidence
106Muse Glimmer 30B43.6%measured
107Ornith-1.0-9B43.5%estimated ± 3.8 pp, high confidence
108Qwen3.5-35B-A3B43.4%estimated ± 3.8 pp, high confidence
109Qwen3.5 Flash43.3%estimated ± 2.9 pp, low confidence
110LongCat-Flash-Lite-Sparse42.9%estimated ± 3.8 pp, high confidence
111Qwen3.6-27B42.7%estimated ± 3.2 pp, high confidence
112GPT-5.6 Sol41.9%estimated ± 2.9 pp, low confidence
113Claude Sonnet 541.5%estimated ± 2.9 pp, low confidence
114Ling 3.0 Flash41.2%measured
115Hy3 Preview41.2%measured
116GLM-4.741.1%estimated ± 2.9 pp, low confidence
117MiniMax M341.1%estimated ± 2.9 pp, low confidence
118A.X K241.0%measured
119Claude Sonnet 4.640.9%estimated ± 2.9 pp, low confidence
120Kimi K2.7 Code40.9%estimated ± 2.9 pp, low confidence
121Trinity-Large-Preview40.6%estimated ± 3.2 pp, high confidence
122Trinity-Large-Thinking40.6%estimated ± 3.2 pp, high confidence
123Ling 3.0 Flash FP840.4%measured
124o140.4%estimated ± 4.7 pp, high confidence
125Mistral Medium 3.5 128B40.2%estimated ± 3.2 pp, high confidence
126Gemma 4 26B A4B40.0%estimated ± 3.2 pp, high confidence
127Qwen3.5-122B-A10B39.7%estimated ± 3.2 pp, high confidence
128GPT-5.4 mini39.7%estimated ± 2.9 pp, low confidence
129MiMo-V2.539.7%estimated ± 2.9 pp, low confidence
130GLM-5.139.5%estimated ± 2.9 pp, low confidence
131MiMo-V2.5-Pro39.5%estimated ± 2.9 pp, low confidence
132GPT-5 (high)39.5%estimated ± 4.7 pp, high confidence
133Ternary Bonsai 2 27B39.1%estimated ± 3.8 pp, high confidence
134DeepSeek V3 032439.1%estimated ± 3.2 pp, high confidence
135GPT-OSS 20B39.0%estimated ± 3.2 pp, high confidence
136Mistral Small 438.9%estimated ± 3.2 pp, high confidence
137Mistral Small 4 (Reasoning)38.9%estimated ± 3.2 pp, high confidence
138North Mini Code38.9%estimated ± 3.2 pp, high confidence
139Granite 4.2 30B38.8%measured
140GLM-4.638.8%estimated ± 2.9 pp, low confidence
141Command A+38.6%estimated ± 3.2 pp, high confidence
142Mercury 2.538.0%measured
143o1-preview37.9%estimated ± 4.7 pp, high confidence
144GLM-5.337.8%estimated ± 2.9 pp, low confidence
145GLM-5.3-Flash37.8%estimated ± 2.9 pp, low confidence
146K-EXAONE 2.037.4%measured
147K-Exaone37.1%estimated ± 4.7 pp, high confidence
148Gemini 3.1 Flash-Lite37.0%estimated ± 2.9 pp, low confidence
149Mistral Large 336.7%estimated ± 3.2 pp, high confidence
150Qwen3.6-35B-A3B36.7%estimated ± 3.2 pp, high confidence
151MiniMax M2.736.6%estimated ± 2.9 pp, low confidence
152Gemma 4 12B36.6%estimated ± 4.7 pp, high confidence
153Nemotron 3 Super 100B36.3%estimated ± 3.2 pp, high confidence
154Granite 4.2 8B36.1%measured
155DeepSeek V336.0%estimated ± 3.2 pp, high confidence
156GPT-4.135.9%estimated ± 3.8 pp, high confidence
157ZAYA1-74B-Preview35.1%estimated ± 3.8 pp, high confidence
158Gemini 3.5 Flash-Lite34.9%estimated ± 2.9 pp, low confidence
159GPT-OSS 120B34.2%estimated ± 3.2 pp, high confidence
160Gemini 1.5 Pro33.4%estimated ± 4.7 pp, high confidence
161o3-mini33.0%estimated ± 3.8 pp, high confidence
162LLaDA2.2-flash33.0%estimated ± 3.8 pp, high confidence
163Claude 3.5 Sonnet32.8%estimated ± 3.8 pp, high confidence
164GPT-4 Turbo32.4%estimated ± 4.7 pp, high confidence
165Nemotron 3 Nano Omni 30B A3B32.0%measured
166Llama 4 Maverick32.0%estimated ± 3.2 pp, high confidence
167Claude 3 Opus31.6%estimated ± 4.7 pp, high confidence
168Nemotron 3.5 Lightning 30B A3B NVFP431.4%measured
169Nemotron 3 Nano 30B30.9%estimated ± 3.2 pp, high confidence
170Mellum2-12B-A2.5B-Thinking30.4%estimated ± 5.4 pp, medium confidence
171Agents-A1-4B29.6%measured
172Ultravox v0.6 Llama 3.3 70B28.3%estimated ± 4.7 pp, medium confidence
173GPT-4o mini28.1%estimated ± 4.7 pp, medium confidence
174GPT-4.1 nano28.0%estimated ± 4.7 pp, medium confidence
175ZAYA1-8B27.2%estimated ± 5.4 pp, medium confidence
176Ling 2.6 Flash27.0%measured
177MiniCPM5-2B26.3%measured
178Gemma 4 E2B26.3%estimated ± 4.7 pp, medium confidence
179Solar Pro 326.0%estimated ± 3.2 pp, high confidence
180Gemma 4 E4B24.9%estimated ± 3.2 pp, medium confidence
181Ling 3.0 Tiny24.7%estimated ± 3.2 pp, medium confidence
182Granite 4.2 3B24.1%measured
183Gemma 3 27B23.9%estimated ± 3.2 pp, medium confidence
184Celeris-122.2%estimated ± 3.2 pp, medium confidence
185Llama 4 Scout21.9%estimated ± 3.2 pp, medium confidence
186GLM-5.219.2%estimated ± 2.9 pp, low confidence
187Laguna M.117.3%estimated ± 2.9 pp, low confidence
188GPT-4.1 mini17.3%estimated ± 3.8 pp, medium confidence
189Laguna XS.216.9%estimated ± 2.9 pp, low confidence
190GLM-4.516.4%estimated ± 2.9 pp, low confidence
191LFM2.5-2.6B15.2%estimated ± 3.2 pp, medium confidence
192Mellum2-12B-A2.5B-Instruct11.4%estimated ± 5.4 pp, low confidence
193Grok Code Fast 110.5%estimated ± 2.9 pp, low confidence
194MiniCPM5-1B10.0%estimated ± 5.4 pp, low confidence
195LLaDA2.2-mini8.0%estimated ± 5.4 pp, low confidence
196Claude Haiku 4.51.0%estimated ± 2.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General