benchgap
Knowledge & reasoning

ARC-AGI-3 leaderboard

As of 2026-10-07, the highest measured score on ARC-AGI-3 is 62.7% by GPT-6 Astra. 222 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.173.5%estimated ± 3.4 pp, medium confidence
2Claude Opus 5.573.3%estimated ± 3.4 pp, medium confidence
3Claude Fable 573.1%estimated ± 3.4 pp, medium confidence
4GPT-6 Astra62.7%measured
5GPT-6.1 Sol52.7%measured
6Sakana Fugu34.2%estimated ± 2.1 pp, high confidence
7Sakana Fugu-Ultra34.2%estimated ± 2.1 pp, high confidence
8Claude Opus 530.2%measured
9Gemini 3.8 Flash10.4%measured
10GPT-5.6 Sol7.8%measured
11GPT-5.4 Pro5.6%estimated ± 11.5 pp, low confidence
12GPT-6 Sol4.6%measured
13Claude Mythos 52.8%estimated ± 2.1 pp, high confidence
14Grok 4.62.1%measured
15Gemini 3 Pro1.8%estimated ± 3.4 pp, high confidence
16Gemini 3.7 Flash1.7%estimated ± 3.4 pp, high confidence
17Claude Sonnet 5.51.7%estimated ± 3.4 pp, high confidence
18GPT-5.3 Codex1.7%estimated ± 3.4 pp, high confidence
19Muse Spark 1.11.7%estimated ± 3.4 pp, high confidence
20A.X K21.7%estimated ± 3.4 pp, medium confidence
21Apodex 1.11.7%estimated ± 3.4 pp, medium confidence
22Apodex 1.1 Mini1.7%estimated ± 3.4 pp, medium confidence
23Celeris-11.7%estimated ± 3.4 pp, medium confidence
24Claude 3 Haiku1.7%estimated ± 3.4 pp, medium confidence
25Claude 4 Sonnet1.7%estimated ± 3.4 pp, medium confidence
26Claude Opus 4.5 Thinking1.7%estimated ± 3.4 pp, high confidence
27Claude Opus 4.6 (Adaptive)1.7%estimated ± 3.4 pp, high confidence
28Claude Opus 4.71.7%estimated ± 3.4 pp, high confidence
29Claude Sonnet 51.7%estimated ± 3.4 pp, medium confidence
30Command A+1.7%estimated ± 3.4 pp, medium confidence
31DeepSeek-R11.7%estimated ± 3.4 pp, medium confidence
32DeepSeek V3 03241.7%estimated ± 3.4 pp, medium confidence
33DeepSeek V3.11.7%estimated ± 3.4 pp, medium confidence
34DeepSeek V3.1 (Reasoning)1.7%estimated ± 3.4 pp, medium confidence
35DeepSeek V3.21.7%estimated ± 3.4 pp, medium confidence
36Exaone 4.0 1.2B1.7%estimated ± 3.4 pp, medium confidence
37Exaone 4.0 32B1.7%estimated ± 3.4 pp, medium confidence
38Gemini 2.5 Flash1.7%estimated ± 3.4 pp, medium confidence
39Gemini 3.5 Flash-Lite1.7%estimated ± 3.4 pp, medium confidence
40Gemini 3.6 Flash1.7%estimated ± 3.4 pp, high confidence
41Gemini 3 Flash1.7%estimated ± 3.4 pp, high confidence
42Gemini 4 Argon1.7%estimated ± 3.4 pp, high confidence
43Gemma 3 27B1.7%estimated ± 3.4 pp, medium confidence
44Gemma 4 26B A4B1.7%estimated ± 3.4 pp, medium confidence
45GLM-4.5-Air1.7%estimated ± 3.4 pp, medium confidence
46GLM-4.61.7%estimated ± 3.4 pp, medium confidence
47GLM-5.11.7%estimated ± 3.4 pp, medium confidence
48GLM-5.31.7%estimated ± 3.4 pp, medium confidence
49GLM-5-Turbo1.7%estimated ± 3.4 pp, medium confidence
50GLM-5V-Turbo1.7%estimated ± 3.4 pp, medium confidence
51GPT-4o1.7%estimated ± 3.4 pp, medium confidence
52GPT-5.11.7%estimated ± 3.4 pp, medium confidence
53GPT-5.1-Codex1.7%estimated ± 3.4 pp, medium confidence
54GPT-5.1-Codex-Max1.7%estimated ± 3.4 pp, medium confidence
55GPT-5.2-Codex1.7%estimated ± 3.4 pp, medium confidence
56GPT-5 (high)1.7%estimated ± 3.4 pp, medium confidence
57GPT-5 (medium)1.7%estimated ± 3.4 pp, medium confidence
58GPT-OSS 120B1.7%estimated ± 3.4 pp, medium confidence
59GPT-OSS 20B1.7%estimated ± 3.4 pp, medium confidence
60Granite-4.0-350M1.7%estimated ± 3.4 pp, medium confidence
61Granite-4.0-H-1B1.7%estimated ± 3.4 pp, medium confidence
62Granite-4.0-H-350M1.7%estimated ± 3.4 pp, medium confidence
63Grok 41.7%estimated ± 3.4 pp, medium confidence
64Grok 4.1 Fast1.7%estimated ± 3.4 pp, medium confidence
65Grok 4.1 Fast (Reasoning)1.7%estimated ± 3.4 pp, medium confidence
66Grok 4.71.7%estimated ± 3.4 pp, high confidence
67Grok 4 Fast (Reasoning)1.7%estimated ± 3.4 pp, medium confidence
68Grok Code Fast 11.7%estimated ± 3.4 pp, medium confidence
69Hy31.7%estimated ± 3.4 pp, medium confidence
70K-Exaone1.7%estimated ± 3.4 pp, medium confidence
71K-EXAONE 2.01.7%estimated ± 3.4 pp, medium confidence
72Kimi K21.7%estimated ± 3.4 pp, medium confidence
73Kimi K2.7 Code1.7%estimated ± 3.4 pp, medium confidence
74LFM2.5-2.6B1.7%estimated ± 3.4 pp, medium confidence
75LFM2.5-8B-A1B1.7%estimated ± 3.4 pp, medium confidence
76LFM2.5-VL-1.6B-Extract1.7%estimated ± 3.4 pp, medium confidence
77Ling 3.0 Flash VL1.7%estimated ± 3.4 pp, medium confidence
78Ling 3.0 Tiny1.7%estimated ± 3.4 pp, medium confidence
79Ling 3.1 Flash1.7%estimated ± 3.4 pp, medium confidence
80Llama 3.1 405B1.7%estimated ± 3.4 pp, medium confidence
81Llama 4 Maverick1.7%estimated ± 3.4 pp, medium confidence
82Llama 4 Scout1.7%estimated ± 3.4 pp, medium confidence
83Mercury 2.51.7%estimated ± 3.4 pp, medium confidence
84MiMo-V2.5-Pro1.7%estimated ± 3.4 pp, medium confidence
85MiMo-V2.6-Flash1.7%estimated ± 3.4 pp, medium confidence
86MiMo-V2.6-Pro1.7%estimated ± 3.4 pp, medium confidence
87MiMo-V2-Omni1.7%estimated ± 3.4 pp, medium confidence
88MiMo-V2-Pro1.7%estimated ± 3.4 pp, medium confidence
89MiniCPM5-2B1.7%estimated ± 3.4 pp, medium confidence
90MiniMax M2.71.7%estimated ± 3.4 pp, medium confidence
91MiniMax M31.7%estimated ± 3.4 pp, medium confidence
92Mistral Large 21.7%estimated ± 3.4 pp, medium confidence
93Mistral Large 31.7%estimated ± 3.4 pp, medium confidence
94Mistral Large 41.7%estimated ± 3.4 pp, medium confidence
95Mistral Medium 31.7%estimated ± 3.4 pp, medium confidence
96Mistral Medium 3.5 128B1.7%estimated ± 3.4 pp, medium confidence
97Mistral Small 41.7%estimated ± 3.4 pp, medium confidence
98Mistral Small 4 (Reasoning)1.7%estimated ± 3.4 pp, medium confidence
99Muse Glimmer 30B1.7%estimated ± 3.4 pp, medium confidence
100Muse Spark1.7%estimated ± 3.4 pp, high confidence
101Muse Spark 1.21.7%estimated ± 3.4 pp, high confidence
102Muse Spark 1.31.7%estimated ± 3.4 pp, high confidence
103Nemotron 3 Nano 30B1.7%estimated ± 3.4 pp, medium confidence
104Nemotron 3 Super 100B1.7%estimated ± 3.4 pp, medium confidence
105Nemotron Ultra 253B1.7%estimated ± 3.4 pp, medium confidence
106North Mini Code1.7%estimated ± 3.4 pp, medium confidence
107Nova Pro1.7%estimated ± 3.4 pp, medium confidence
108o31.7%estimated ± 3.4 pp, medium confidence
109Phi-41.7%estimated ± 3.4 pp, medium confidence
110Quasar 438B1.7%estimated ± 3.4 pp, medium confidence
111Qwen3.5 397B (Reasoning)1.7%estimated ± 3.4 pp, medium confidence
112Qwen 3.6 Max (preview)1.7%estimated ± 3.4 pp, medium confidence
113Qwen3.8 Max Preview1.7%estimated ± 3.4 pp, medium confidence
114Qwen3 Max1.7%estimated ± 3.4 pp, medium confidence
115Qwen3-Omni-30B-A3B-Instruct1.7%estimated ± 3.4 pp, medium confidence
116Qwen3-Omni-30B-A3B-Thinking1.7%estimated ± 3.4 pp, medium confidence
117Sarvam 105B1.7%estimated ± 3.4 pp, medium confidence
118Sarvam 30B1.7%estimated ± 3.4 pp, medium confidence
119Solar Pro 21.7%estimated ± 3.4 pp, medium confidence
120Solar Pro 31.7%estimated ± 3.4 pp, medium confidence
121Solar Pro 41.7%estimated ± 3.4 pp, medium confidence
122Step 3.7 Flash1.7%estimated ± 3.4 pp, medium confidence
123Step 5 Preview1.7%estimated ± 3.4 pp, medium confidence
124Trinity-Large-Preview1.7%estimated ± 3.4 pp, medium confidence
125Trinity-Large-Thinking1.7%estimated ± 3.4 pp, medium confidence
126Ultravox v0.6 Llama 3.3 70B1.7%estimated ± 3.4 pp, medium confidence
127Claude Haiku 5.51.6%estimated ± 10.7 pp, low confidence
128Claude 3 Opus1.6%estimated ± 10.7 pp, low confidence
129Claude 4.1 Opus Thinking1.6%estimated ± 10.7 pp, low confidence
130DeepSeek R1 Distill Qwen 32B1.6%estimated ± 10.7 pp, low confidence
131Gemini 1.0 Pro1.6%estimated ± 10.7 pp, low confidence
132Gemini 1.5 Pro1.6%estimated ± 10.7 pp, low confidence
133GPT-4 Turbo1.6%estimated ± 10.7 pp, low confidence
134GPT-4o mini1.6%estimated ± 10.7 pp, low confidence
135Phi-4 Multimodal Instruct1.6%estimated ± 10.7 pp, low confidence
136Qwen2.5 Coder 32B Instruct1.6%estimated ± 10.7 pp, low confidence
137Claude Opus 4.81.5%measured
138Kimi K30.9%estimated ± 2.1 pp, high confidence
139GPT-5.6 Terra0.8%measured
140GPT-5.5 Pro0.6%estimated ± 11.5 pp, low confidence
141Pareto 26.90.5%estimated ± 11.6 pp, low confidence
142GPT-5.50.4%measured
143Gemini 3.1 Pro0.4%measured
144Grok 4.50.3%measured
145Agents-A10.3%estimated ± 11.5 pp, low confidence
146dots3-note Preview0.3%estimated ± 11.5 pp, low confidence
147Ornith-1.5-397B0.2%estimated ± 2.1 pp, high confidence
148GPT-5.40.2%measured
149Claude Opus 4.7 (Adaptive)0.2%measured
150GPT-5.6 Luna0.2%measured
151Qwen3.8 Max0.1%estimated ± 2.1 pp, high confidence
152GPT-6 Luna0.1%measured
153GPT-5.20.1%estimated ± 2.1 pp, high confidence
154Qwen3.7 Max0.1%estimated ± 2.1 pp, high confidence
155Pareto 26.10 Preview0.1%estimated ± 3.8 pp, high confidence
156Grok 4.200.1%measured
157Hy4 preview0.1%estimated ± 2.1 pp, high confidence
158Gemini 3.5 Flash0.1%estimated ± 2.1 pp, medium confidence
159Qwen3.8-Flash-Next0.0%estimated ± 2.1 pp, medium confidence
160Claude Opus 4.60.0%estimated ± 2.1 pp, medium confidence
161GLM-5.20.0%estimated ± 2.1 pp, medium confidence
162Qwen3.8-Omni-Flash0.0%estimated ± 2.1 pp, medium confidence
163DeepSeek V4.1 Flash0.0%estimated ± 2.1 pp, medium confidence
164Kimi K2.60.0%estimated ± 2.1 pp, medium confidence
165Beam0.0%estimated ± 3.8 pp, high confidence
166Qwen3.6 Plus0.0%estimated ± 2.1 pp, medium confidence
167Qwen3.7 Plus0.0%estimated ± 2.1 pp, medium confidence
168DeepSeek V4 Pro 08130.0%estimated ± 2.1 pp, medium confidence
169Grok 4.30.0%estimated ± 2.1 pp, medium confidence
170Claude Sonnet 4.60.0%estimated ± 2.1 pp, medium confidence
171Interfaze Beta0.0%estimated ± 2.1 pp, medium confidence
172Inkling-Small0.0%estimated ± 2.1 pp, medium confidence
173Ornith-1.5-35B-A3B0.0%estimated ± 2.1 pp, medium confidence
174Qwen3.8-27B0.0%estimated ± 2.1 pp, medium confidence
175Claude 3.5 Sonnet0.0%estimated ± 2.1 pp, medium confidence
176Claude Haiku 4.50.0%estimated ± 9.2 pp, low confidence
177Claude Opus 4.50.0%estimated ± 2.1 pp, medium confidence
178Claude Sonnet 4.50.0%estimated ± 2.1 pp, medium confidence
179Claude Sonnet 4.5 Thinking0.0%estimated ± 13.8 pp, low confidence
180DeepSeek V30.0%estimated ± 2.1 pp, medium confidence
181DeepSeek V4 Flash 07310.0%estimated ± 2.1 pp, medium confidence
182Gemini 2.5 Pro0.0%estimated ± 2.1 pp, medium confidence
183Gemini 3.1 Flash-Lite0.0%estimated ± 9.2 pp, low confidence
184Gemini 3 Pro Deep Think0.0%estimated ± 13.8 pp, low confidence
185Gemma 4 12B0.0%estimated ± 2.1 pp, medium confidence
186Gemma 4 31B0.0%estimated ± 2.1 pp, medium confidence
187Gemma 4 E2B0.0%estimated ± 2.1 pp, medium confidence
188Gemma 4 E4B0.0%estimated ± 2.1 pp, medium confidence
189GLM-4.50.0%estimated ± 9.2 pp, low confidence
190GLM-4.70.0%estimated ± 2.1 pp, medium confidence
191GLM-50.0%estimated ± 2.1 pp, medium confidence
192GLM-5.3-Flash0.0%estimated ± 9.2 pp, low confidence
193GPT-4.10.0%estimated ± 2.1 pp, medium confidence
194GPT-4.1 mini0.0%estimated ± 2.1 pp, medium confidence
195GPT-4.1 nano0.0%estimated ± 2.1 pp, medium confidence
196GPT-5.4 mini0.0%estimated ± 2.1 pp, medium confidence
197GPT-5.4 nano0.0%estimated ± 2.1 pp, medium confidence
198Granite 4.2 30B0.0%estimated ± 2.1 pp, medium confidence
199Granite 4.2 3B0.0%estimated ± 2.1 pp, medium confidence
200Granite 4.2 8B0.0%estimated ± 2.1 pp, medium confidence
201Hy3 Preview0.0%estimated ± 2.1 pp, medium confidence
202Inkling0.0%estimated ± 2.1 pp, medium confidence
203Kimi K2.50.0%estimated ± 2.1 pp, medium confidence
204Kimi K2.5 (Reasoning)0.0%estimated ± 2.1 pp, medium confidence
205Laguna M.10.0%estimated ± 9.2 pp, low confidence
206Laguna XS.20.0%estimated ± 9.2 pp, low confidence
207LFM2.5-230M0.0%estimated ± 2.1 pp, medium confidence
208LFM2.5-VL-450M0.0%estimated ± 2.1 pp, medium confidence
209Ling 2.6 Flash0.0%estimated ± 2.1 pp, medium confidence
210Ling 3.0 Flash0.0%estimated ± 2.1 pp, medium confidence
211Ling 3.0 Flash FP80.0%estimated ± 2.1 pp, medium confidence
212LLaDA2.2-mini0.0%estimated ± 3.8 pp, medium confidence
213LongCat-Flash-Lite-Sparse0.0%estimated ± 3.8 pp, medium confidence
214MAI-Thinking-10.0%estimated ± 2.1 pp, medium confidence
215Mellum2-12B-A2.5B-Instruct0.0%estimated ± 2.1 pp, medium confidence
216Mellum2-12B-A2.5B-Thinking0.0%estimated ± 2.1 pp, medium confidence
217MiMo-V2.50.0%estimated ± 9.2 pp, low confidence
218MiMo-V2-Flash0.0%estimated ± 2.1 pp, medium confidence
219MiniCPM5-1B0.0%estimated ± 3.8 pp, medium confidence
220Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 2.1 pp, medium confidence
221Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 2.1 pp, medium confidence
222Nemotron 3 Ultra0.0%estimated ± 2.1 pp, medium confidence
223o10.0%estimated ± 2.1 pp, medium confidence
224o1-pro0.0%estimated ± 2.1 pp, medium confidence
225o3-mini0.0%estimated ± 2.1 pp, medium confidence
226Ornith-1.5-9B0.0%estimated ± 2.1 pp, medium confidence
227Qwen3 235B 25070.0%estimated ± 2.1 pp, medium confidence
228Qwen3.5-122B-A10B0.0%estimated ± 2.1 pp, medium confidence
229Qwen3.5-27B0.0%estimated ± 2.1 pp, medium confidence
230Qwen3.5-35B-A3B0.0%estimated ± 2.1 pp, medium confidence
231Qwen3.5 397B0.0%estimated ± 2.1 pp, medium confidence
232Qwen3.5 Flash0.0%estimated ± 9.2 pp, low confidence
233Qwen3.6-27B0.0%estimated ± 2.1 pp, medium confidence
234Qwen3.6-35B-A3B0.0%estimated ± 2.1 pp, medium confidence
235Solar Open 20.0%estimated ± 3.8 pp, medium confidence
236Soofi S 30B-A3B0.0%estimated ± 2.1 pp, medium confidence
237Ternary Bonsai 2 27B0.0%estimated ± 2.1 pp, medium confidence
238ZAYA1-74B-Preview0.0%estimated ± 2.1 pp, medium confidence
239ZAYA1-8B0.0%estimated ± 2.1 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General