benchgap
Coding

AA-SciCode leaderboard

As of 2026-10-07, the highest measured score on AA-SciCode is 66.9% by Claude Opus 5.5. 112 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.566.9%measured
2Claude Fable 5.163.1%measured
3Gemini 4 Argon61.8%measured
4Claude Fable 561.0%measured
5Claude Sonnet 5.561.0%measured
6Claude Mythos 561.0%estimated ± 3.8 pp, high confidence
7MiMo-V2.6-Pro60.9%measured
8Kimi K359.5%measured
9GLM-5.359.0%measured
10Step 5 Preview58.9%measured
11Sakana Fugu58.8%estimated ± 3.1 pp, medium confidence
12Muse Spark 1.158.8%measured
13Muse Spark 1.358.8%measured
14Gemini 3.1 Pro58.7%measured
15SWE-258.5%estimated ± 3.8 pp, medium confidence
16GPT-6 Sol57.6%measured
17Composer 2.557.6%estimated ± 1.7 pp, medium confidence
18Sakana Fugu-Ultra57.5%estimated ± 3.1 pp, high confidence
19Grok 4.757.4%measured
20Muse Spark 1.257.4%measured
21Gemini 3.7 Flash57.2%measured
22GPT-5.6 Sol57.1%measured
23Composer 256.6%estimated ± 8.6 pp, low confidence
24Gemini 3.8 Flash56.6%measured
25GPT-6 Astra56.5%measured
26Grok 4.656.5%measured
27Claude Opus 556.4%measured
28Composer 2 Fast56.1%estimated ± 8.6 pp, low confidence
29SWE-1.756.0%estimated ± 3.8 pp, medium confidence
30Claude Opus 4.755.8%estimated ± 3.5 pp, high confidence
31GPT-5.555.8%measured
32Claude Opus 4.7 (Adaptive)55.4%estimated ± 3.2 pp, high confidence
33Claude Haiku 5.555.0%measured
34GPT-5.6 Terra55.0%measured
35Grok 4.555.0%measured
36GPT-6 Luna54.6%measured
37Claude Opus 4.854.4%measured
38Claude Sonnet 554.3%measured
39GPT-5.454.3%estimated ± 3.2 pp, high confidence
40GPT-6.1 Sol54.2%measured
41Mistral Large 454.2%measured
42Ling 3.1 Flash54.1%measured
43Gemini 3.5 Flash53.9%measured
44Ember-153.6%estimated ± 3.9 pp, high confidence
45GPT-5.6 Luna53.6%measured
46Claude Opus 4.6 (Adaptive)53.5%estimated ± 3.8 pp, high confidence
47Gemini 3.6 Flash53.4%measured
48Ornith-1.5-397B52.8%estimated ± 3.0 pp, high confidence
49Hy4 preview52.7%estimated ± 3.0 pp, high confidence
50Qwen3.8 Max52.2%estimated ± 3.0 pp, high confidence
51Claude Sonnet 4.652.2%estimated ± 1.7 pp, medium confidence
52GPT-5.4 mini52.1%measured
53Qwen3.8 Max Preview52.1%measured
54DeepSeek V4.1 Flash51.9%measured
55GLM-5.3-Flash51.6%measured
56Kimi K2.651.5%measured
57MiMo-V2.6-Flash51.3%measured
58GPT-5.2-Codex51.2%estimated ± 3.8 pp, high confidence
59GLM-5.251.2%measured
60DeepSeek V4 Pro 081351.0%measured
61dots3-note Preview50.8%estimated ± 3.0 pp, high confidence
62MiMo-V2.5-Pro50.6%measured
63Qwen3.8-Flash-Next50.6%measured
64Qwen3.8-Omni-Flash50.5%estimated ± 3.0 pp, high confidence
65DeepSeek V4 Flash 073150.3%measured
66Ornith-1.0-397B50.2%estimated ± 3.0 pp, high confidence
67Atria Dawn Preview50.1%estimated ± 3.8 pp, high confidence
68MiniMax M2.750.1%measured
69Laguna S 2.150.0%estimated ± 3.8 pp, high confidence
70Seed 2.1 Pro49.8%estimated ± 3.0 pp, high confidence
71Inkling-Small49.7%measured
72GLM-4.549.6%estimated ± 4.6 pp, low confidence
73Qwen3.7 Max49.5%measured
74Ornith-1.5-35B-A3B49.5%estimated ± 3.0 pp, high confidence
75Beam49.2%estimated ± 3.1 pp, high confidence
76GLM-5 (Reasoning)49.1%estimated ± 3.8 pp, high confidence
77Claude Sonnet 4.5 Thinking49.0%estimated ± 3.8 pp, high confidence
78GPT-5.1-Codex-Max49.0%estimated ± 3.8 pp, high confidence
79Claude Opus 4.5 Thinking48.7%estimated ± 3.8 pp, high confidence
80Muse Spark48.7%estimated ± 3.2 pp, high confidence
81GPT-5.3 Codex48.7%estimated ± 3.8 pp, high confidence
82Gemini 3 Flash48.7%estimated ± 3.8 pp, high confidence
83Solar Open 248.6%estimated ± 3.7 pp, medium confidence
84Hy348.6%measured
85Hy3 Preview48.6%measured
86MiMo-V2.548.3%estimated ± 3.8 pp, high confidence
87Grok 4.348.3%measured
88Seed 2.1 Turbo48.3%estimated ± 3.0 pp, high confidence
89Kimi K2.548.3%estimated ± 3.1 pp, high confidence
90Quasar 438B48.1%measured
91GPT-5.248.1%estimated ± 3.8 pp, high confidence
92Claude Opus 4.548.0%estimated ± 3.0 pp, high confidence
93Qwen3.5 Plus48.0%estimated ± 3.8 pp, high confidence
94MiniMax M2.547.9%estimated ± 3.8 pp, high confidence
95Qwen 3.6 Max (preview)47.9%estimated ± 3.0 pp, high confidence
96Gemini 3 Pro47.8%estimated ± 3.8 pp, high confidence
97Kimi K2.7 Code47.8%measured
98GPT-5 mini47.8%estimated ± 3.8 pp, high confidence
99GLM-547.8%estimated ± 3.8 pp, high confidence
100GPT-5.1-Codex47.6%estimated ± 3.8 pp, high confidence
101Claude Haiku 4.5 Thinking47.4%estimated ± 3.8 pp, high confidence
102MiMo-V2-Pro47.3%estimated ± 3.9 pp, high confidence
103Ternary Bonsai 2 27B47.2%estimated ± 3.7 pp, high confidence
104GPT-5.4 nano47.2%measured
105MiniMax M347.1%measured
106Inkling47.0%measured
107Claude Sonnet 4.546.9%estimated ± 3.9 pp, high confidence
108Claude Opus 4.646.9%estimated ± 3.8 pp, high confidence
109Qwen3.6 Plus46.9%estimated ± 3.2 pp, high confidence
110Qwen3.8-27B46.6%measured
111DeepSeek V3.2 (Thinking)46.5%estimated ± 3.8 pp, high confidence
112Gemini 2.5 Pro46.3%measured
113Qwen3 Max46.2%estimated ± 3.8 pp, high confidence
114GLM-4.646.2%estimated ± 3.8 pp, high confidence
115Qwen3.7 Plus46.1%measured
116Grok 4.2046.1%estimated ± 3.8 pp, high confidence
117Grok 446.0%estimated ± 8.6 pp, low confidence
118Grok 4.1 Fast (Reasoning)45.9%estimated ± 3.8 pp, high confidence
119BTL-345.9%estimated ± 3.7 pp, high confidence
120MiMo-V2-Omni45.8%estimated ± 3.9 pp, high confidence
121Gemini 3.1 Flash-Lite45.7%estimated ± 3.8 pp, medium confidence
122Grok 4 Fast (Reasoning)45.7%estimated ± 3.8 pp, medium confidence
123Claude 4.1 Opus45.7%estimated ± 3.9 pp, high confidence
124MAI-Thinking-145.6%estimated ± 3.7 pp, high confidence
125Apodex 1.145.5%measured
126Apodex 1.1 Mini45.5%measured
127Gemma 4 31B45.5%measured
128DeepSeek V3.245.4%estimated ± 8.6 pp, low confidence
129Qwen3.5 Flash45.2%estimated ± 4.3 pp, medium confidence
130Claude Haiku 4.545.1%estimated ± 3.9 pp, high confidence
131Muse Glimmer 30B44.9%measured
132Claude 4 Sonnet44.8%estimated ± 3.9 pp, high confidence
133GLM-5.144.8%measured
134MiMo-V2-Flash44.8%estimated ± 3.2 pp, high confidence
135MAI-Code-1.1-Flash44.8%estimated ± 3.9 pp, high confidence
136Laguna M.144.7%estimated ± 3.8 pp, high confidence
137Qwen3.5-27B44.7%estimated ± 3.9 pp, high confidence
138Solar Pro 444.6%measured
139GPT-5.144.6%estimated ± 3.2 pp, high confidence
140Ling 3.0 Flash VL44.2%measured
141Step 3.7 Flash43.9%measured
142Grok Code Fast 143.9%estimated ± 3.9 pp, high confidence
143Laguna XS 2.143.9%estimated ± 3.8 pp, high confidence
144Kimi K2.5 (Reasoning)43.4%estimated ± 3.2 pp, high confidence
145Laguna XS.243.2%estimated ± 3.8 pp, high confidence
146Qwen3.5-35B-A3B43.1%estimated ± 3.9 pp, high confidence
147Qwen3.6-27B42.8%measured
148GLM-4.742.7%estimated ± 3.2 pp, high confidence
149Qwen3.5 397B42.3%estimated ± 3.7 pp, high confidence
150K-EXAONE 2.042.0%measured
151Ling 3.0 Flash42.0%measured
152Ling 3.0 Flash FP842.0%measured
153Ornith-1.0-35B41.8%estimated ± 3.0 pp, high confidence
154Gemini 3.5 Flash-Lite41.3%measured
155A.X K241.0%measured
156Trinity-Large-Preview40.6%measured
157Trinity-Large-Thinking40.6%measured
158Nemotron 3 Ultra40.3%measured
159o140.2%estimated ± 3.2 pp, high confidence
160LongCat-Flash-Lite-Sparse40.2%estimated ± 3.8 pp, high confidence
161Mistral Medium 3.5 128B40.2%measured
162Gemma 4 26B A4B40.0%measured
163Qwen3.5-122B-A10B39.7%measured
164Ornith-1.5-9B39.7%estimated ± 3.0 pp, high confidence
165GPT-5 (high)39.4%estimated ± 3.2 pp, high confidence
166DeepSeek V3 032439.0%measured
167GPT-OSS 20B38.9%measured
168Mistral Small 438.8%measured
169Mistral Small 4 (Reasoning)38.8%measured
170North Mini Code38.8%measured
171Command A+38.5%measured
172Mercury 2.538.5%measured
173Granite 4.2 30B37.8%measured
174o1-preview37.7%estimated ± 3.2 pp, high confidence
175K-Exaone36.8%estimated ± 3.2 pp, high confidence
176Mistral Large 336.6%measured
177Qwen3.6-35B-A3B36.6%measured
178Gemma 4 12B36.3%estimated ± 3.2 pp, high confidence
179Nemotron 3 Super 100B36.2%measured
180DeepSeek V335.8%measured
181GPT-4.135.6%estimated ± 3.9 pp, high confidence
182LLaDA2.2-flash34.7%estimated ± 3.8 pp, high confidence
183Ornith-1.0-9B34.6%estimated ± 3.0 pp, medium confidence
184GPT-OSS 120B34.0%measured
185Gemini 1.5 Pro33.1%estimated ± 3.2 pp, high confidence
186Nemotron 3 Nano Omni 30B A3B32.9%estimated ± 3.1 pp, high confidence
187o3-mini32.8%estimated ± 3.9 pp, high confidence
188Claude 3.5 Sonnet32.6%estimated ± 3.9 pp, high confidence
189Nemotron 3.5 Lightning 30B A3B NVFP432.1%measured
190GPT-4 Turbo32.1%estimated ± 3.2 pp, high confidence
191Llama 4 Maverick31.7%measured
192GPT-4.1 mini31.5%estimated ± 3.2 pp, high confidence
193Granite 4.2 8B31.5%measured
194Claude 3 Opus31.2%estimated ± 3.2 pp, high confidence
195Agents-A1-4B30.6%estimated ± 3.1 pp, high confidence
196Nemotron 3 Nano 30B30.6%measured
197Ling 2.6 Flash28.2%estimated ± 3.1 pp, high confidence
198Ultravox v0.6 Llama 3.3 70B27.8%estimated ± 3.2 pp, high confidence
199GPT-4o mini27.6%estimated ± 3.2 pp, high confidence
200GPT-4.1 nano27.5%estimated ± 3.2 pp, high confidence
201Mellum2-12B-A2.5B-Thinking27.3%estimated ± 3.7 pp, high confidence
202MiniCPM5-2B26.3%measured
203Gemma 4 E2B25.7%estimated ± 3.2 pp, medium confidence
204Solar Pro 325.5%measured
205Granite 4.2 3B25.3%measured
206Gemma 4 E4B24.4%measured
207Ling 3.0 Tiny24.2%measured
208Gemma 3 27B23.3%measured
209BTL-422.5%estimated ± 3.7 pp, high confidence
210ZAYA1-8B22.1%estimated ± 3.7 pp, high confidence
211ZAYA1-74B-Preview21.9%estimated ± 3.7 pp, high confidence
212Celeris-121.6%measured
213Llama 4 Scout21.3%measured
214LFM2.5-2.6B14.4%measured
215Mellum2-12B-A2.5B-Instruct1.1%estimated ± 3.7 pp, medium confidence
216MiniCPM5-1B0.6%estimated ± 3.7 pp, medium confidence
217LLaDA2.2-mini0.2%estimated ± 3.7 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General