benchgap
Knowledge & reasoning

ARC-AGI-2 leaderboard

As of 2026-10-07, the highest measured score on ARC-AGI-2 is 95.0% by GPT-6 Astra. 200 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra95.0%measured
2GPT-6.1 Sol94.2%measured
3GPT-5.6 Sol92.5%measured
4Claude Opus 5.591.7%measured
5Sakana Fugu91.5%estimated ± 10.1 pp, low confidence
6Sakana Fugu-Ultra91.5%estimated ± 10.1 pp, low confidence
7Claude Opus 590.4%measured
8Claude Fable 5.190.0%measured
9GPT-6 Sol89.6%measured
10Claude Fable 589.2%measured
11Gemini 3.8 Flash89.2%measured
12GPT-5.585.0%measured
13Claude Sonnet 5.584.8%estimated ± 13.0 pp, low confidence
14Gemini 3.7 Flash84.6%measured
15GPT-5.5 Pro84.2%measured
16GPT-5.6 Terra83.9%measured
17GPT-5.4 Pro83.3%measured
18Gemini 4 Argon83.0%estimated ± 14.2 pp, low confidence
19dots3-note Preview81.4%measured
20Pareto 26.980.9%estimated ± 13.0 pp, low confidence
21Muse Spark 1.380.9%estimated ± 14.2 pp, low confidence
22Grok 4.779.8%estimated ± 14.2 pp, low confidence
23MiMo-V2.6-Pro79.7%estimated ± 14.2 pp, low confidence
24Claude Mythos 579.5%estimated ± 11.6 pp, low confidence
25Qwen3.8 Max Preview79.1%estimated ± 14.2 pp, low confidence
26Claude Haiku 5.578.0%estimated ± 13.0 pp, low confidence
27Gemini 3.1 Pro77.1%measured
28Claude Opus 4.7 (Adaptive)75.8%measured
29Ling 3.1 Flash74.9%estimated ± 14.2 pp, low confidence
30Step 5 Preview74.7%estimated ± 10.1 pp, low confidence
31GPT-5.474.0%measured
32MiniMax M372.7%estimated ± 12.6 pp, low confidence
33Gemini 3.5 Flash72.1%measured
34Claude Opus 4.872.1%measured
35Mistral Large 471.1%estimated ± 14.2 pp, low confidence
36MiMo-V2.6-Flash70.3%estimated ± 14.2 pp, low confidence
37Ornith-1.5-397B70.1%estimated ± 10.1 pp, low confidence
38Qwen3.8 Max68.9%estimated ± 10.1 pp, low confidence
39Claude Opus 4.6 (Adaptive)68.8%measured
40Pareto 26.10 Preview67.7%estimated ± 10.1 pp, low confidence
41Qwen3.7 Max67.7%estimated ± 10.1 pp, low confidence
42Grok 4.667.1%measured
43Hy4 preview67.1%estimated ± 10.1 pp, low confidence
44Muse Spark 1.266.8%estimated ± 13.0 pp, low confidence
45Muse Spark 1.166.2%estimated ± 12.6 pp, low confidence
46Qwen3.8-Flash-Next63.8%estimated ± 10.1 pp, low confidence
47Claude Sonnet 562.5%estimated ± 3.1 pp, low confidence
48Claude Opus 4.761.8%estimated ± 12.6 pp, low confidence
49DeepSeek V4 Flash 073161.4%measured
50DeepSeek V4 Pro 081361.3%measured
51GLM-5.261.2%estimated ± 10.1 pp, low confidence
52Gemini 3.6 Flash60.4%measured
53Kimi K360.4%measured
54Qwen3.8-Omni-Flash60.3%estimated ± 10.1 pp, low confidence
55DeepSeek V4.1 Flash59.8%estimated ± 10.1 pp, low confidence
56GPT-5.6 Luna59.5%measured
57GPT-6 Luna59.3%measured
58GPT-5.3 Codex58.9%estimated ± 14.2 pp, low confidence
59Claude Sonnet 4.658.3%measured
60Kimi K2.657.9%estimated ± 10.1 pp, low confidence
61Beam57.9%estimated ± 10.1 pp, low confidence
62Qwen3.7 Plus57.0%estimated ± 10.1 pp, low confidence
63Qwen3.6 Plus55.6%estimated ± 11.6 pp, low confidence
64Interfaze Beta55.3%estimated ± 10.1 pp, low confidence
65Grok 4.353.6%estimated ± 11.6 pp, low confidence
66Grok 4.2053.3%measured
67GPT-5.252.9%measured
68GLM-5.352.7%estimated ± 12.6 pp, low confidence
69Grok 4.552.6%measured
70Claude Opus 4.652.6%estimated ± 10.1 pp, low confidence
71Ornith-1.5-35B-A3B52.6%estimated ± 10.1 pp, low confidence
72Qwen3.8-27B52.6%estimated ± 10.1 pp, low confidence
73Gemini 3 Flash51.9%estimated ± 12.6 pp, low confidence
74Solar Pro 451.8%estimated ± 10.1 pp, low confidence
75Inkling48.0%estimated ± 10.1 pp, low confidence
76MiMo-V2-Pro47.1%estimated ± 14.2 pp, low confidence
77Kimi K2.547.0%estimated ± 10.1 pp, low confidence
78GPT-5.2-Codex46.6%estimated ± 14.2 pp, low confidence
79Qwen 3.6 Max (preview)46.2%estimated ± 14.2 pp, low confidence
80Hy3 Preview45.8%estimated ± 10.1 pp, low confidence
81GLM-5.3-Flash45.4%estimated ± 12.6 pp, low confidence
82MiniMax M2.745.2%estimated ± 10.1 pp, low confidence
83Nemotron 3 Ultra45.2%estimated ± 10.1 pp, low confidence
84Gemini 3 Pro Deep Think45.1%measured
85Ornith-1.5-9B43.5%estimated ± 10.1 pp, low confidence
86Solar Open 243.2%estimated ± 10.1 pp, low confidence
87GLM-5.142.9%estimated ± 10.1 pp, low confidence
88Qwen3.5 397B42.7%estimated ± 11.6 pp, low confidence
89Muse Spark42.5%measured
90GLM-542.4%estimated ± 10.1 pp, low confidence
91Ternary Bonsai 2 27B41.8%estimated ± 10.1 pp, low confidence
92A.X K241.4%estimated ± 10.1 pp, low confidence
93Quasar 438B40.4%estimated ± 14.2 pp, low confidence
94Inkling-Small40.1%measured
95GLM-5-Turbo39.9%estimated ± 14.2 pp, low confidence
96Ling 3.0 Flash39.8%estimated ± 10.1 pp, low confidence
97Apodex 1.139.2%estimated ± 14.2 pp, low confidence
98Apodex 1.1 Mini39.2%estimated ± 14.2 pp, low confidence
99Qwen3.6-27B38.8%estimated ± 11.6 pp, low confidence
100MAI-Thinking-138.0%estimated ± 10.1 pp, low confidence
101Claude Opus 4.5 Thinking37.6%measured
102Ling 3.0 Flash FP837.6%estimated ± 10.1 pp, low confidence
103Kimi K2.5 (Reasoning)37.5%estimated ± 11.6 pp, low confidence
104Kimi K2.7 Code37.0%estimated ± 14.2 pp, low confidence
105Hy335.1%estimated ± 14.2 pp, low confidence
106Gemini 3.5 Flash-Lite34.1%estimated ± 12.6 pp, low confidence
107K-EXAONE 2.033.9%estimated ± 10.1 pp, low confidence
108Claude Opus 4.533.6%estimated ± 11.6 pp, low confidence
109Ling 3.0 Flash VL32.4%estimated ± 14.2 pp, low confidence
110Gemini 3 Pro31.1%measured
111Qwen3.5-122B-A10B31.0%estimated ± 11.6 pp, low confidence
112MiMo-V2-Omni30.0%estimated ± 14.2 pp, low confidence
113Qwen3.5 Flash29.8%estimated ± 12.6 pp, low confidence
114GPT-5.1-Codex29.2%estimated ± 14.2 pp, low confidence
115GPT-5.1-Codex-Max29.2%estimated ± 14.2 pp, low confidence
116MiMo-V2.5-Pro28.9%estimated ± 12.6 pp, low confidence
117Mercury 2.528.7%estimated ± 10.1 pp, low confidence
118GLM-5V-Turbo28.5%estimated ± 14.2 pp, low confidence
119Gemma 4 12B28.4%estimated ± 10.1 pp, low confidence
120GPT-5.128.3%estimated ± 5.7 pp, medium confidence
121Qwen3.6-35B-A3B27.2%estimated ± 11.6 pp, low confidence
122GPT-5 (high)26.6%estimated ± 14.2 pp, low confidence
123GPT-5 (medium)26.2%estimated ± 14.2 pp, low confidence
124Claude 4.1 Opus Thinking26.2%estimated ± 14.2 pp, low confidence
125GLM-4.725.2%estimated ± 11.6 pp, low confidence
126Trinity-Large-Thinking25.1%estimated ± 10.1 pp, low confidence
127Command A+24.9%estimated ± 14.2 pp, low confidence
128Grok 424.9%estimated ± 14.2 pp, low confidence
129MiMo-V2.524.6%estimated ± 12.6 pp, low confidence
130Nemotron 3.5 Lightning 30B A3B NVFP424.3%estimated ± 10.1 pp, low confidence
131Qwen3.5-27B23.9%estimated ± 11.6 pp, low confidence
132o3-pro22.7%estimated ± 14.2 pp, low confidence
133Gemini 3.1 Flash-Lite22.4%estimated ± 12.6 pp, low confidence
134Qwen3.5 397B (Reasoning)21.3%estimated ± 14.2 pp, low confidence
135Nemotron 3 Nano Omni 30B A3B20.9%estimated ± 10.1 pp, low confidence
136ZAYA1-8B19.8%estimated ± 10.1 pp, low confidence
137MiniCPM5-2B19.1%estimated ± 10.1 pp, low confidence
138GPT-5.4 mini18.9%measured
139LongCat-Flash-Lite-Sparse18.6%estimated ± 10.1 pp, low confidence
140Grok 4.1 Fast (Reasoning)17.8%estimated ± 14.2 pp, low confidence
141o317.3%estimated ± 14.2 pp, low confidence
142Gemma 4 31B16.2%estimated ± 11.6 pp, low confidence
143Qwen3.5-35B-A3B15.5%estimated ± 11.6 pp, low confidence
144Step 3.7 Flash15.2%estimated ± 14.2 pp, low confidence
145Trinity-Large-Preview14.4%estimated ± 10.1 pp, low confidence
146Claude Sonnet 4.5 Thinking13.6%measured
147Claude Sonnet 4.513.6%measured
148Claude 4.1 Opus12.7%estimated ± 14.2 pp, low confidence
149MiMo-V2-Flash12.3%estimated ± 11.6 pp, low confidence
150Mellum2-12B-A2.5B-Thinking11.6%estimated ± 10.1 pp, low confidence
151ZAYA1-74B-Preview11.5%estimated ± 10.1 pp, low confidence
152Grok 4 Fast (Reasoning)11.1%estimated ± 14.2 pp, low confidence
153Muse Glimmer 30B10.0%estimated ± 14.2 pp, low confidence
154Claude 4 Sonnet8.2%estimated ± 14.2 pp, low confidence
155Gemini 2.5 Pro7.8%estimated ± 11.6 pp, low confidence
156DeepSeek V3.27.1%estimated ± 14.2 pp, low confidence
157LLaDA2.2-mini7.0%estimated ± 10.1 pp, low confidence
158Soofi S 30B-A3B6.7%estimated ± 10.1 pp, low confidence
159Qwen3 Max6.4%estimated ± 14.2 pp, low confidence
160Mellum2-12B-A2.5B-Instruct6.1%estimated ± 10.1 pp, low confidence
161GPT-5.4 nano5.7%measured
162K-Exaone4.5%estimated ± 14.2 pp, low confidence
163Grok Code Fast 14.1%estimated ± 14.2 pp, low confidence
164DeepSeek V3.13.7%estimated ± 14.2 pp, low confidence
165DeepSeek V3.1 (Reasoning)3.4%estimated ± 14.2 pp, low confidence
166MiniCPM5-1B3.2%estimated ± 10.1 pp, low confidence
167LFM2.5-230M3.1%estimated ± 10.1 pp, low confidence
168DeepSeek-R13.1%estimated ± 14.2 pp, low confidence
169Nemotron 3 Super 100B2.8%estimated ± 14.2 pp, low confidence
170Kimi K22.7%estimated ± 14.2 pp, low confidence
171GPT-OSS 120B1.8%estimated ± 14.2 pp, low confidence
172o1-preview1.6%estimated ± 14.2 pp, low confidence
173Grok 4.1 Fast1.6%estimated ± 14.2 pp, low confidence
174Mistral Small 41.6%estimated ± 14.2 pp, low confidence
175Mistral Small 4 (Reasoning)1.6%estimated ± 14.2 pp, low confidence
176GLM-4.5-Air1.5%estimated ± 14.2 pp, low confidence
177Ling 3.0 Tiny1.5%estimated ± 14.2 pp, low confidence
178Llama 4 Maverick0.9%estimated ± 14.2 pp, low confidence
179North Mini Code0.9%estimated ± 14.2 pp, low confidence
180Gemini 2.5 Flash0.9%estimated ± 14.2 pp, low confidence
181DeepSeek V3 03240.8%estimated ± 14.2 pp, low confidence
182Mistral Large 30.7%estimated ± 14.2 pp, low confidence
183Mistral Medium 30.6%estimated ± 14.2 pp, low confidence
184GPT-OSS 20B0.6%estimated ± 14.2 pp, low confidence
185Nemotron 3 Nano 30B0.6%estimated ± 14.2 pp, low confidence
186Sarvam 105B0.5%estimated ± 14.2 pp, low confidence
187Claude 3 Opus0.5%estimated ± 14.2 pp, low confidence
188GPT-4o0.4%estimated ± 14.2 pp, low confidence
189LFM2.5-2.6B0.4%estimated ± 14.2 pp, low confidence
190DeepSeek R1 Distill Qwen 32B0.4%estimated ± 14.2 pp, low confidence
191Llama 4 Scout0.4%estimated ± 14.2 pp, low confidence
192Gemini 1.5 Pro0.3%estimated ± 14.2 pp, low confidence
193Solar Pro 30.3%estimated ± 14.2 pp, low confidence
194Qwen3-Omni-30B-A3B-Thinking0.3%estimated ± 14.2 pp, low confidence
195Ultravox v0.6 Llama 3.3 70B0.3%estimated ± 14.2 pp, low confidence
196Mistral Large 20.3%estimated ± 14.2 pp, low confidence
197Nemotron Ultra 253B0.3%estimated ± 14.2 pp, low confidence
198Llama 3.1 405B0.2%estimated ± 14.2 pp, low confidence
199LFM2.5-8B-A1B0.2%estimated ± 14.2 pp, low confidence
200GPT-4 Turbo0.2%estimated ± 14.2 pp, low confidence
201Solar Pro 20.2%estimated ± 14.2 pp, low confidence
202Nova Pro0.2%estimated ± 14.2 pp, low confidence
203Qwen2.5 Coder 32B Instruct0.2%estimated ± 14.2 pp, low confidence
204GPT-4o mini0.2%estimated ± 14.2 pp, low confidence
205Sarvam 30B0.1%estimated ± 14.2 pp, low confidence
206Celeris-10.1%estimated ± 14.2 pp, low confidence
207Exaone 4.0 32B0.1%estimated ± 14.2 pp, low confidence
208Qwen3-Omni-30B-A3B-Instruct0.1%estimated ± 14.2 pp, low confidence
209Phi-40.1%estimated ± 14.2 pp, low confidence
210Phi-4 Multimodal Instruct0.1%estimated ± 14.2 pp, low confidence
211Claude 3 Haiku0.1%estimated ± 14.2 pp, low confidence
212Gemini 1.0 Pro0.1%estimated ± 14.2 pp, low confidence
213Exaone 4.0 1.2B0.1%estimated ± 14.2 pp, low confidence
214Granite-4.0-H-1B0.1%estimated ± 14.2 pp, low confidence
215Gemma 3 27B0.0%estimated ± 14.2 pp, low confidence
216Granite-4.0-350M0.0%estimated ± 14.2 pp, low confidence
217Granite-4.0-H-350M0.0%estimated ± 14.2 pp, low confidence
218LFM2.5-VL-1.6B-Extract0.0%estimated ± 14.2 pp, low confidence
219Gemma 4 26B A4B0.0%estimated ± 13.0 pp, low confidence
220Claude 3.5 Sonnet0.0%estimated ± 11.6 pp, low confidence
221Claude Haiku 4.50.0%estimated ± 12.6 pp, low confidence
222DeepSeek V30.0%estimated ± 11.6 pp, low confidence
223Gemma 4 E2B0.0%estimated ± 11.6 pp, low confidence
224Gemma 4 E4B0.0%estimated ± 11.6 pp, low confidence
225GLM-4.50.0%estimated ± 12.6 pp, low confidence
226GLM-4.60.0%estimated ± 12.6 pp, low confidence
227GPT-4.10.0%estimated ± 11.6 pp, low confidence
228GPT-4.1 mini0.0%estimated ± 11.6 pp, low confidence
229GPT-4.1 nano0.0%estimated ± 11.6 pp, low confidence
230Granite 4.2 30B0.0%estimated ± 11.6 pp, low confidence
231Granite 4.2 3B0.0%estimated ± 11.6 pp, low confidence
232Granite 4.2 8B0.0%estimated ± 11.6 pp, low confidence
233Laguna M.10.0%estimated ± 12.6 pp, low confidence
234Laguna XS.20.0%estimated ± 12.6 pp, low confidence
235LFM2.5-VL-450M0.0%estimated ± 11.6 pp, low confidence
236Ling 2.6 Flash0.0%estimated ± 11.6 pp, low confidence
237Mistral Medium 3.5 128B0.0%estimated ± 12.6 pp, low confidence
238o10.0%estimated ± 11.6 pp, low confidence
239o1-pro0.0%estimated ± 11.6 pp, low confidence
240o3-mini0.0%estimated ± 11.6 pp, low confidence
241Qwen3 235B 25070.0%estimated ± 11.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General