benchgap
Coding

AA Coding Index leaderboard

As of 2026-10-07, the highest measured score on AA Coding Index is 81.6% by Claude Fable 5.1. 117 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Composer 288.9%estimated ± 11.4 pp, low confidence
2Composer 2 Fast87.1%estimated ± 11.4 pp, low confidence
3Claude Opus 5.583.9%estimated ± 1.0 pp, medium confidence
4Claude Sonnet 5.583.0%estimated ± 1.0 pp, medium confidence
5Claude Fable 5.181.6%measured
6Grok 4.779.0%estimated ± 1.0 pp, high confidence
7Claude Mythos 578.9%estimated ± 5.1 pp, medium confidence
8Gemini 4 Argon78.6%estimated ± 2.5 pp, medium confidence
9Claude Opus 578.0%measured
10Ember-177.8%estimated ± 2.5 pp, high confidence
11DeepSeek V4.1 Flash77.5%estimated ± 2.5 pp, high confidence
12Pareto 26.977.4%estimated ± 2.5 pp, high confidence
13GPT-5.6 Sol77.4%measured
14SWE-277.1%estimated ± 2.5 pp, high confidence
15GPT-6 Astra76.9%measured
16Grok 4.676.8%measured
17GPT-6.1 Sol76.7%estimated ± 2.5 pp, high confidence
18MiMo-V2.6-Pro76.7%estimated ± 2.5 pp, high confidence
19GPT-5.6 Terra76.7%measured
20Claude Fable 576.5%measured
21Gemini 3.8 Flash76.3%measured
22Kimi K376.2%measured
23Pareto 26.10 Preview76.1%estimated ± 2.5 pp, high confidence
24Gemini 3.7 Flash76.1%measured
25GPT-6 Sol75.8%estimated ± 2.5 pp, high confidence
26Muse Spark 1.375.8%measured
27MiMo-V2.6-Flash75.5%estimated ± 2.5 pp, high confidence
28Step 5 Preview75.5%estimated ± 2.5 pp, high confidence
29Claude Haiku 5.575.2%estimated ± 3.3 pp, high confidence
30GPT-6 Luna75.1%estimated ± 2.5 pp, high confidence
31GPT-5.574.9%measured
32GLM-5.374.8%measured
33Hy4 preview74.4%estimated ± 2.5 pp, high confidence
34Claude Opus 4.874.3%measured
35GLM-5.3-Flash74.1%estimated ± 2.5 pp, high confidence
36Claude Opus 4.7 (Adaptive)73.6%measured
37Claude Opus 4.773.5%estimated ± 3.1 pp, high confidence
38SWE-1.773.5%estimated ± 3.3 pp, high confidence
39Qwen3.8-Flash-Next73.1%measured
40Grok 4.572.5%measured
41Qwen3.8-Omni-Flash72.4%estimated ± 2.5 pp, high confidence
42Muse Spark 1.272.2%measured
43Qwen3.8 Max72.0%estimated ± 2.5 pp, high confidence
44Ornith-1.5-397B71.8%estimated ± 2.5 pp, high confidence
45Qwen3.8 Max Preview71.8%measured
46Claude Sonnet 571.6%measured
47GPT-5.6 Luna71.5%measured
48Muse Spark 1.171.3%measured
49GPT-5.471.1%measured
50Mistral Large 470.6%estimated ± 6.4 pp, medium confidence
51Ling 3.1 Flash70.4%estimated ± 6.4 pp, medium confidence
52Gemini 3.5 Flash70.1%measured
53Gemini 3.6 Flash69.2%measured
54DeepSeek V4 Flash 073169.1%measured
55DeepSeek V4 Pro 081368.8%measured
56Gemini 3.1 Pro68.8%measured
57GLM-5.268.8%measured
58GPT-5.3 Codex68.6%estimated ± 5.1 pp, medium confidence
59Beam68.3%estimated ± 2.5 pp, high confidence
60dots3-note Preview68.1%estimated ± 4.6 pp, high confidence
61Qwen3.8-27B68.1%measured
62Claude Opus 4.6 (Adaptive)67.8%estimated ± 6.8 pp, medium confidence
63Sakana Fugu-Ultra67.5%estimated ± 6.4 pp, low confidence
64Laguna S 2.167.0%estimated ± 2.5 pp, medium confidence
65Ornith-1.0-397B66.9%estimated ± 4.6 pp, high confidence
66Sakana Fugu66.9%estimated ± 6.4 pp, low confidence
67Qwen3.7 Max66.0%measured
68Composer 2.565.9%estimated ± 1.0 pp, medium confidence
69Seed 2.1 Pro65.8%estimated ± 4.6 pp, high confidence
70Atria Dawn Preview64.8%estimated ± 7.5 pp, medium confidence
71Seed 2.1 Turbo62.3%estimated ± 4.6 pp, high confidence
72Kimi K2.661.8%measured
73Claude Opus 4.561.7%estimated ± 4.6 pp, high confidence
74Ornith-1.5-35B-A3B61.4%estimated ± 2.5 pp, medium confidence
75Qwen 3.6 Max (preview)61.3%estimated ± 4.6 pp, high confidence
76Quasar 438B61.2%measured
77Apodex 1.160.8%measured
78Apodex 1.1 Mini60.8%measured
79Kimi K2.7 Code60.8%measured
80GPT-5.260.2%estimated ± 5.1 pp, medium confidence
81MiMo-V2.5-Pro60.2%measured
82GLM-4.559.6%estimated ± 9.0 pp, low confidence
83Gemini 3 Flash59.4%estimated ± 5.8 pp, medium confidence
84Hy358.8%measured
85Hy3 Preview58.8%measured
86Muse Spark58.6%measured
87MiniMax M358.6%measured
88BTL-357.4%estimated ± 6.4 pp, medium confidence
89BTL-457.2%estimated ± 5.1 pp, medium confidence
90MiMo-V2-Pro56.4%estimated ± 5.1 pp, medium confidence
91GPT-5.2-Codex56.2%estimated ± 5.8 pp, medium confidence
92GPT-5.4 mini56.1%measured
93GPT-5.4 nano56.1%measured
94GLM-556.0%estimated ± 5.1 pp, medium confidence
95Claude Opus 4.655.9%estimated ± 3.3 pp, medium confidence
96Qwen3.7 Plus55.9%measured
97GLM-5.155.8%measured
98Claude Sonnet 4.554.8%estimated ± 5.1 pp, medium confidence
99Qwen3.6 Plus54.5%measured
100MiMo-V2.554.3%estimated ± 5.8 pp, medium confidence
101Grok 4.2053.8%estimated ± 5.1 pp, medium confidence
102Qwen3.6-27B53.7%measured
103Inkling-Small52.9%measured
104Qwen3.5 397B52.8%estimated ± 5.1 pp, medium confidence
105MiniMax M2.752.6%measured
106Inkling52.1%measured
107GLM-5 (Reasoning)51.4%estimated ± 6.8 pp, medium confidence
108Claude Sonnet 4.5 Thinking51.0%estimated ± 6.8 pp, medium confidence
109GPT-5.1-Codex-Max50.7%estimated ± 6.8 pp, medium confidence
110Ling 3.0 Flash50.6%measured
111Ling 3.0 Flash FP850.6%measured
112MiMo-V2-Omni50.0%estimated ± 5.1 pp, medium confidence
113Claude Opus 4.5 Thinking49.9%estimated ± 6.8 pp, medium confidence
114MiMo-V2-Flash49.8%measured
115Laguna M.149.6%estimated ± 5.1 pp, medium confidence
116Claude Sonnet 4.649.5%estimated ± 3.3 pp, medium confidence
117Claude 4.1 Opus49.4%estimated ± 5.1 pp, medium confidence
118GPT-5.149.4%measured
119Gemini 3.5 Flash-Lite49.3%measured
120Nemotron 3 Ultra49.3%measured
121Muse Glimmer 30B49.0%measured
122Ornith-1.0-35B48.8%estimated ± 4.6 pp, high confidence
123Ling 3.0 Flash VL48.7%estimated ± 6.4 pp, medium confidence
124MAI-Thinking-147.5%estimated ± 5.1 pp, medium confidence
125Qwen3.5 Plus47.2%estimated ± 6.8 pp, medium confidence
126Claude Haiku 4.547.1%estimated ± 5.1 pp, medium confidence
127Mistral Medium 3.5 128B46.9%measured
128Kimi K2.546.8%measured
129Kimi K2.5 (Reasoning)46.8%measured
130MiniMax M2.546.7%estimated ± 6.8 pp, medium confidence
131Gemini 3 Pro46.4%estimated ± 6.8 pp, medium confidence
132GPT-5 mini46.4%estimated ± 6.8 pp, medium confidence
133Claude 4 Sonnet46.0%estimated ± 5.1 pp, medium confidence
134GPT-5.1-Codex45.8%estimated ± 6.8 pp, medium confidence
135MAI-Code-1.1-Flash45.8%estimated ± 5.1 pp, medium confidence
136Qwen3.5-122B-A10B45.7%measured
137Ornith-1.5-9B45.5%estimated ± 4.6 pp, high confidence
138Qwen3.5-27B45.4%estimated ± 5.1 pp, medium confidence
139Qwen3.5 Flash45.3%estimated ± 5.8 pp, medium confidence
140GLM-4.745.3%measured
141Claude Haiku 4.5 Thinking44.9%estimated ± 6.8 pp, medium confidence
142Gemma 4 31B43.4%measured
143Gemini 3.1 Flash-Lite43.1%estimated ± 5.8 pp, medium confidence
144Laguna XS 2.142.7%estimated ± 5.1 pp, medium confidence
145Grok Code Fast 142.5%estimated ± 5.1 pp, medium confidence
146Grok 4.342.3%measured
147Solar Pro 442.1%estimated ± 5.1 pp, medium confidence
148Qwen3.6-35B-A3B41.9%measured
149Solar Open 241.8%estimated ± 5.1 pp, medium confidence
150Grok 441.6%estimated ± 11.4 pp, low confidence
151DeepSeek V3.2 (Thinking)41.4%estimated ± 6.8 pp, medium confidence
152Laguna XS.240.9%estimated ± 5.1 pp, medium confidence
153Qwen3 Max40.5%estimated ± 6.8 pp, medium confidence
154GLM-4.640.3%estimated ± 6.8 pp, medium confidence
155A.X K240.3%estimated ± 6.4 pp, medium confidence
156Qwen3.5-35B-A3B39.7%estimated ± 5.1 pp, medium confidence
157o139.7%measured
158Ornith-1.0-9B39.6%estimated ± 4.6 pp, medium confidence
159Step 3.7 Flash39.6%measured
160Gemma 4 26B A4B39.3%measured
161Grok 4.1 Fast (Reasoning)39.3%estimated ± 6.8 pp, medium confidence
162DeepSeek V3.239.1%estimated ± 11.4 pp, low confidence
163Grok 4 Fast (Reasoning)38.6%estimated ± 6.8 pp, low confidence
164K-EXAONE 2.038.2%estimated ± 5.1 pp, medium confidence
165LongCat-Flash-Lite-Sparse38.2%estimated ± 5.1 pp, medium confidence
166GPT-5 (high)37.8%measured
167Nemotron 3 Super 100B37.7%measured
168DeepSeek V3 032435.2%estimated ± 6.4 pp, medium confidence
169North Mini Code34.7%estimated ± 6.4 pp, medium confidence
170o1-preview34.1%measured
171Mercury 2.534.0%estimated ± 6.4 pp, medium confidence
172Gemini 2.5 Pro33.3%measured
173K-Exaone32.1%measured
174Gemma 4 12B31.0%measured
175GPT-OSS 120B30.4%measured
176Ternary Bonsai 2 27B29.4%estimated ± 5.1 pp, medium confidence
177Command A+27.9%measured
178Nemotron 3.5 Lightning 30B A3B NVFP426.8%measured
179Granite 4.2 30B26.7%estimated ± 5.1 pp, medium confidence
180Mistral Small 426.6%measured
181Mistral Small 4 (Reasoning)26.6%measured
182Trinity-Large-Preview25.8%measured
183Trinity-Large-Thinking25.8%measured
184GPT-4.125.5%estimated ± 5.1 pp, medium confidence
185Ling 2.6 Flash25.3%measured
186ZAYA1-74B-Preview24.9%estimated ± 5.1 pp, medium confidence
187o3-mini23.8%estimated ± 5.1 pp, medium confidence
188LLaDA2.2-flash23.8%estimated ± 5.1 pp, medium confidence
189Claude 3.5 Sonnet23.7%estimated ± 5.1 pp, medium confidence
190Gemini 1.5 Pro23.6%measured
191MiniCPM5-2B23.2%estimated ± 5.1 pp, medium confidence
192DeepSeek V323.0%measured
193Granite 4.2 8B22.4%measured
194Mellum2-12B-A2.5B-Thinking22.3%estimated ± 6.4 pp, medium confidence
195GPT-4 Turbo21.5%measured
196GPT-OSS 20B20.7%measured
197GPT-4.1 mini20.2%measured
198Mistral Large 320.1%measured
199Claude 3 Opus19.5%measured
200ZAYA1-8B16.5%estimated ± 6.4 pp, medium confidence
201Llama 4 Maverick16.3%measured
202Celeris-114.4%measured
203Nemotron 3 Nano 30B14.4%measured
204Nemotron 3 Nano Omni 30B A3B13.8%measured
205Solar Pro 312.9%estimated ± 6.4 pp, medium confidence
206Granite 4.2 3B12.7%estimated ± 6.4 pp, medium confidence
207Ultravox v0.6 Llama 3.3 70B11.9%measured
208Ling 3.0 Tiny11.8%estimated ± 6.4 pp, medium confidence
209GPT-4o mini11.4%measured
210GPT-4.1 nano11.1%measured
211Gemma 3 27B10.1%measured
212Agents-A1-4B9.7%estimated ± 6.4 pp, medium confidence
213Gemma 4 E4B9.4%measured
214Llama 4 Scout8.2%measured
215LFM2.5-2.6B7.7%measured
216Gemma 4 E2B7.2%measured
217Mellum2-12B-A2.5B-Instruct0.6%estimated ± 6.4 pp, low confidence
218MiniCPM5-1B0.3%estimated ± 6.4 pp, low confidence
219LLaDA2.2-mini0.1%estimated ± 6.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General