benchgap
Knowledge & reasoning

ARC-AGI-1 leaderboard

As of 2026-10-07, the highest measured score on ARC-AGI-1 is 98.5% by Claude Fable 5. 207 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 4 Argon99.7%estimated ± 0.7 pp, low confidence
2Claude Fable 598.5%measured
3Gemini 3.8 Flash98.5%measured
4GPT-6 Astra98.5%measured
5Claude Fable 5.197.5%measured
6Claude Opus 597.5%measured
7Claude Opus 5.597.5%measured
8Claude Mythos 597.4%estimated ± 3.7 pp, high confidence
9Claude Sonnet 5.597.2%estimated ± 3.7 pp, high confidence
10Muse Spark 1.196.8%estimated ± 3.7 pp, high confidence
11Sakana Fugu96.5%estimated ± 2.7 pp, high confidence
12Sakana Fugu-Ultra96.5%estimated ± 2.7 pp, high confidence
13GPT-5.6 Sol96.5%measured
14GPT-5.6 Terra96.5%measured
15GPT-6.1 Sol96.5%measured
16Pareto 26.996.3%estimated ± 3.7 pp, high confidence
17Claude Haiku 5.595.5%estimated ± 3.7 pp, high confidence
18Gemini 3.7 Flash95.5%measured
19GPT-6 Sol95.5%measured
20Claude Sonnet 595.4%estimated ± 0.7 pp, low confidence
21dots3-note Preview95.0%estimated ± 1.9 pp, high confidence
22GPT-5.595.0%measured
23GPT-5.5 Pro95.0%measured
24GPT-5.4 Pro94.5%measured
25Kimi K394.5%measured
26Muse Spark 1.394.4%estimated ± 5.4 pp, medium confidence
27Grok 4.794.3%estimated ± 5.4 pp, medium confidence
28MiMo-V2.6-Pro94.2%estimated ± 5.4 pp, medium confidence
29Qwen3.8 Max Preview94.1%estimated ± 5.4 pp, medium confidence
30GLM-5.394.1%estimated ± 5.4 pp, medium confidence
31Step 5 Preview94.1%estimated ± 2.7 pp, high confidence
32Gemini 3.1 Pro94.0%estimated ± 1.9 pp, high confidence
33GPT-5.493.7%measured
34Claude Opus 4.7 (Adaptive)93.7%estimated ± 1.9 pp, high confidence
35GLM-5.3-Flash93.5%estimated ± 5.4 pp, medium confidence
36Ling 3.1 Flash93.4%estimated ± 5.4 pp, medium confidence
37Ornith-1.5-397B93.1%estimated ± 2.7 pp, high confidence
38Claude Opus 4.6 (Adaptive)93.0%measured
39Muse Spark 1.293.0%estimated ± 5.4 pp, medium confidence
40Qwen3.8 Max92.9%estimated ± 2.7 pp, high confidence
41Mistral Large 492.6%estimated ± 5.4 pp, medium confidence
42Pareto 26.10 Preview92.6%estimated ± 2.7 pp, high confidence
43Qwen3.7 Max92.6%estimated ± 2.7 pp, high confidence
44Claude Opus 4.892.5%measured
45Gemini 3.5 Flash92.5%measured
46Hy4 preview92.4%estimated ± 2.7 pp, high confidence
47MiMo-V2.6-Flash92.4%estimated ± 5.4 pp, medium confidence
48Grok 4.392.3%estimated ± 5.4 pp, medium confidence
49Qwen3.8-Flash-Next91.6%estimated ± 2.7 pp, high confidence
50Gemini 3.6 Flash91.2%measured
51GLM-5.290.9%estimated ± 2.7 pp, high confidence
52Qwen3.8-Omni-Flash90.6%estimated ± 2.7 pp, high confidence
53DeepSeek V4.1 Flash90.4%estimated ± 2.7 pp, high confidence
54DeepSeek V4 Pro 081390.0%measured
55Kimi K2.689.9%estimated ± 2.7 pp, high confidence
56Beam89.9%estimated ± 2.7 pp, high confidence
57Qwen3.7 Plus89.6%estimated ± 2.7 pp, high confidence
58GPT-5.3 Codex89.1%estimated ± 5.4 pp, medium confidence
59DeepSeek V4 Flash 073189.0%measured
60Interfaze Beta88.9%estimated ± 2.7 pp, high confidence
61GPT-5.6 Luna88.0%measured
62Claude Opus 4.687.9%estimated ± 2.7 pp, high confidence
63Ornith-1.5-35B-A3B87.9%estimated ± 2.7 pp, high confidence
64Qwen3.8-27B87.9%estimated ± 2.7 pp, high confidence
65Solar Pro 487.5%estimated ± 2.7 pp, high confidence
66Claude Opus 4.787.4%estimated ± 5.4 pp, medium confidence
67Grok 4.687.0%measured
68GPT-6 Luna86.7%measured
69Grok 4.2086.7%estimated ± 1.9 pp, high confidence
70GPT-5.286.2%measured
71Agents-A186.1%estimated ± 11.9 pp, low confidence
72Claude Sonnet 4.686.0%measured
73Inkling85.7%estimated ± 2.7 pp, medium confidence
74Grok 4.585.7%measured
75MiMo-V2.5-Pro85.3%estimated ± 3.7 pp, high confidence
76Kimi K2.585.2%estimated ± 2.7 pp, medium confidence
77MiniMax M384.9%estimated ± 5.4 pp, medium confidence
78Hy3 Preview84.6%estimated ± 2.7 pp, medium confidence
79MiniMax M2.784.2%estimated ± 2.7 pp, medium confidence
80Nemotron 3 Ultra84.2%estimated ± 2.7 pp, medium confidence
81Inkling-Small84.0%measured
82MiMo-V2-Pro83.9%estimated ± 5.4 pp, medium confidence
83GPT-5.2-Codex83.6%estimated ± 5.4 pp, medium confidence
84Qwen 3.6 Max (preview)83.4%estimated ± 5.4 pp, medium confidence
85Gemini 3 Pro Deep Think83.3%estimated ± 1.9 pp, high confidence
86Ornith-1.5-9B83.2%estimated ± 2.7 pp, medium confidence
87Solar Open 283.0%estimated ± 2.7 pp, medium confidence
88GLM-5.182.8%estimated ± 2.7 pp, medium confidence
89GLM-582.4%estimated ± 2.7 pp, medium confidence
90Muse Spark82.1%estimated ± 1.9 pp, high confidence
91Ternary Bonsai 2 27B82.0%estimated ± 2.7 pp, medium confidence
92A.X K281.7%estimated ± 2.7 pp, medium confidence
93Ling 3.0 Flash80.5%estimated ± 2.7 pp, medium confidence
94Qwen3.6 Plus80.4%estimated ± 5.4 pp, medium confidence
95Claude Opus 4.5 Thinking80.0%measured
96Quasar 438B79.7%estimated ± 5.4 pp, medium confidence
97GLM-5-Turbo79.3%estimated ± 5.4 pp, medium confidence
98MAI-Thinking-179.1%estimated ± 2.7 pp, medium confidence
99Apodex 1.178.8%estimated ± 5.4 pp, medium confidence
100Apodex 1.1 Mini78.8%estimated ± 5.4 pp, medium confidence
101Ling 3.0 Flash FP878.7%estimated ± 2.7 pp, medium confidence
102Kimi K2.7 Code77.1%estimated ± 5.4 pp, medium confidence
103Hy375.5%estimated ± 5.4 pp, medium confidence
104K-EXAONE 2.075.1%estimated ± 2.7 pp, medium confidence
105Qwen3.5 Flash75.1%estimated ± 5.9 pp, medium confidence
106Gemini 3 Pro75.0%measured
107Ling 3.0 Flash VL72.9%estimated ± 5.4 pp, medium confidence
108GPT-5.172.8%measured
109MiMo-V2.572.6%estimated ± 5.9 pp, medium confidence
110Gemini 3.1 Flash-Lite71.5%estimated ± 5.9 pp, medium confidence
111MiMo-V2-Omni70.4%estimated ± 5.4 pp, medium confidence
112GPT-5.1-Codex69.4%estimated ± 5.4 pp, medium confidence
113GPT-5.1-Codex-Max69.4%estimated ± 5.4 pp, medium confidence
114Claude Opus 4.569.4%estimated ± 5.4 pp, medium confidence
115GLM-5V-Turbo68.6%estimated ± 5.4 pp, medium confidence
116Kimi K2.5 (Reasoning)68.4%estimated ± 5.4 pp, medium confidence
117Mercury 2.568.2%estimated ± 2.7 pp, medium confidence
118Gemma 4 12B67.8%estimated ± 2.7 pp, medium confidence
119GPT-5 (high)66.2%estimated ± 5.4 pp, medium confidence
120Qwen3.5-27B65.8%estimated ± 5.4 pp, medium confidence
121GPT-5 (medium)65.7%estimated ± 5.4 pp, medium confidence
122Claude 4.1 Opus Thinking65.6%estimated ± 5.4 pp, medium confidence
123Command A+63.9%estimated ± 5.4 pp, medium confidence
124Grok 463.8%estimated ± 5.4 pp, medium confidence
125GPT-5.4 mini63.7%measured
126Claude Sonnet 4.5 Thinking63.7%measured
127GLM-4.762.6%estimated ± 5.4 pp, medium confidence
128Gemini 3.5 Flash-Lite62.2%estimated ± 5.4 pp, medium confidence
129Trinity-Large-Thinking62.0%estimated ± 2.7 pp, medium confidence
130Claude Sonnet 4.561.4%estimated ± 1.9 pp, high confidence
131o3-pro60.6%estimated ± 5.4 pp, medium confidence
132Nemotron 3.5 Lightning 30B A3B NVFP460.3%estimated ± 2.7 pp, medium confidence
133Qwen3.5 397B58.3%estimated ± 5.4 pp, medium confidence
134Qwen3.5 397B (Reasoning)58.3%estimated ± 5.4 pp, medium confidence
135Qwen3.6-27B58.2%estimated ± 5.4 pp, medium confidence
136Nemotron 3 Nano Omni 30B A3B52.1%estimated ± 2.7 pp, medium confidence
137Grok 4.1 Fast (Reasoning)51.9%estimated ± 5.4 pp, low confidence
138GPT-5.4 nano51.5%measured
139o350.8%estimated ± 5.4 pp, low confidence
140Claude Haiku 4.550.8%estimated ± 5.9 pp, low confidence
141GLM-4.550.8%estimated ± 5.9 pp, low confidence
142ZAYA1-8B49.2%estimated ± 2.7 pp, medium confidence
143MiniCPM5-2B47.2%estimated ± 2.7 pp, medium confidence
144Step 3.7 Flash46.2%estimated ± 5.4 pp, low confidence
145LongCat-Flash-Lite-Sparse45.5%estimated ± 2.7 pp, medium confidence
146Qwen3.5-35B-A3B45.2%estimated ± 5.4 pp, low confidence
147Claude 4.1 Opus40.2%estimated ± 5.4 pp, low confidence
148Qwen3 235B 250739.5%estimated ± 6.9 pp, low confidence
149Qwen3.6-35B-A3B37.9%estimated ± 5.4 pp, low confidence
150Grok 4 Fast (Reasoning)36.0%estimated ± 5.4 pp, low confidence
151Gemini 3 Flash35.9%estimated ± 5.4 pp, low confidence
152Muse Glimmer 30B32.9%estimated ± 5.4 pp, low confidence
153Trinity-Large-Preview31.0%estimated ± 2.7 pp, medium confidence
154Gemma 4 31B28.5%estimated ± 3.7 pp, medium confidence
155Claude 4 Sonnet27.3%estimated ± 5.4 pp, low confidence
156Gemini 2.5 Pro24.1%estimated ± 5.4 pp, low confidence
157MiMo-V2-Flash23.9%estimated ± 5.4 pp, low confidence
158DeepSeek V3.223.8%estimated ± 5.4 pp, low confidence
159Qwen3 Max21.3%estimated ± 5.4 pp, low confidence
160Qwen3.5-122B-A10B21.1%estimated ± 5.4 pp, low confidence
161Mellum2-12B-A2.5B-Thinking19.8%estimated ± 2.7 pp, medium confidence
162ZAYA1-74B-Preview19.3%estimated ± 2.7 pp, medium confidence
163o119.2%estimated ± 5.4 pp, low confidence
164GLM-4.617.6%estimated ± 5.4 pp, low confidence
165Granite 4.2 30B17.1%estimated ± 5.4 pp, low confidence
166Laguna XS.215.2%estimated ± 5.9 pp, low confidence
167K-Exaone14.9%estimated ± 5.4 pp, low confidence
168Gemma 4 26B A4B14.8%estimated ± 3.7 pp, medium confidence
169Mistral Medium 3.5 128B14.1%estimated ± 5.4 pp, low confidence
170Grok Code Fast 113.5%estimated ± 5.4 pp, low confidence
171Ling 2.6 Flash13.4%estimated ± 5.4 pp, low confidence
172DeepSeek V3.112.0%estimated ± 5.4 pp, low confidence
173DeepSeek V3.1 (Reasoning)11.1%estimated ± 5.4 pp, low confidence
174DeepSeek-R19.7%estimated ± 5.4 pp, low confidence
175Nemotron 3 Super 100B8.7%estimated ± 5.4 pp, low confidence
176Kimi K28.4%estimated ± 5.4 pp, low confidence
177GPT-4.18.3%estimated ± 5.4 pp, low confidence
178o3-mini7.6%estimated ± 5.4 pp, low confidence
179o1-pro7.5%estimated ± 5.4 pp, low confidence
180GPT-OSS 120B5.3%estimated ± 5.4 pp, low confidence
181o1-preview4.8%estimated ± 5.4 pp, low confidence
182LLaDA2.2-mini4.8%estimated ± 2.7 pp, medium confidence
183Grok 4.1 Fast4.6%estimated ± 5.4 pp, low confidence
184Mistral Small 44.6%estimated ± 5.4 pp, low confidence
185Mistral Small 4 (Reasoning)4.6%estimated ± 5.4 pp, low confidence
186GLM-4.5-Air4.2%estimated ± 5.4 pp, low confidence
187Granite 4.2 8B4.2%estimated ± 5.4 pp, low confidence
188Soofi S 30B-A3B4.2%estimated ± 2.7 pp, medium confidence
189Ling 3.0 Tiny4.2%estimated ± 5.4 pp, low confidence
190Mellum2-12B-A2.5B-Instruct3.0%estimated ± 2.7 pp, medium confidence
191GPT-4.1 mini2.7%estimated ± 5.4 pp, low confidence
192Llama 4 Maverick2.5%estimated ± 5.4 pp, low confidence
193North Mini Code2.4%estimated ± 5.4 pp, low confidence
194Gemini 2.5 Flash2.3%estimated ± 5.4 pp, low confidence
195DeepSeek V3 03242.1%estimated ± 5.4 pp, low confidence
196Mistral Large 31.7%estimated ± 5.4 pp, low confidence
197Granite 4.2 3B1.5%estimated ± 5.4 pp, low confidence
198Mistral Medium 31.5%estimated ± 5.4 pp, low confidence
199GPT-OSS 20B1.4%estimated ± 5.4 pp, low confidence
200Gemma 4 E4B1.3%estimated ± 5.4 pp, low confidence
201Nemotron 3 Nano 30B1.3%estimated ± 5.4 pp, low confidence
202Sarvam 105B1.3%estimated ± 5.4 pp, low confidence
203Claude 3 Opus1.2%estimated ± 5.4 pp, low confidence
204DeepSeek V31.0%estimated ± 5.4 pp, low confidence
205GPT-4o1.0%estimated ± 5.4 pp, low confidence
206LFM2.5-2.6B1.0%estimated ± 5.4 pp, low confidence
207DeepSeek R1 Distill Qwen 32B1.0%estimated ± 5.4 pp, low confidence
208Llama 4 Scout0.8%estimated ± 5.4 pp, low confidence
209Gemini 1.5 Pro0.7%estimated ± 5.4 pp, low confidence
210GPT-4.1 nano0.7%estimated ± 5.4 pp, low confidence
211Solar Pro 30.7%estimated ± 5.4 pp, low confidence
212Gemma 4 E2B0.7%estimated ± 5.4 pp, low confidence
213Qwen3-Omni-30B-A3B-Thinking0.6%estimated ± 5.4 pp, low confidence
214Ultravox v0.6 Llama 3.3 70B0.6%estimated ± 5.4 pp, low confidence
215Mistral Large 20.6%estimated ± 5.4 pp, low confidence
216Nemotron Ultra 253B0.6%estimated ± 5.4 pp, low confidence
217Llama 3.1 405B0.5%estimated ± 5.4 pp, low confidence
218LFM2.5-8B-A1B0.4%estimated ± 5.4 pp, low confidence
219GPT-4 Turbo0.4%estimated ± 5.4 pp, low confidence
220Solar Pro 20.4%estimated ± 5.4 pp, low confidence
221Nova Pro0.4%estimated ± 5.4 pp, low confidence
222Qwen2.5 Coder 32B Instruct0.3%estimated ± 5.4 pp, low confidence
223GPT-4o mini0.3%estimated ± 5.4 pp, low confidence
224Sarvam 30B0.3%estimated ± 5.4 pp, low confidence
225Laguna M.10.2%estimated ± 5.9 pp, low confidence
226Celeris-10.2%estimated ± 5.4 pp, low confidence
227Exaone 4.0 32B0.2%estimated ± 5.4 pp, low confidence
228MiniCPM5-1B0.2%estimated ± 2.7 pp, medium confidence
229LFM2.5-230M0.2%estimated ± 2.7 pp, medium confidence
230Qwen3-Omni-30B-A3B-Instruct0.2%estimated ± 5.4 pp, low confidence
231Phi-40.2%estimated ± 5.4 pp, low confidence
232Phi-4 Multimodal Instruct0.1%estimated ± 5.4 pp, low confidence
233Claude 3 Haiku0.1%estimated ± 5.4 pp, low confidence
234Gemini 1.0 Pro0.1%estimated ± 5.4 pp, low confidence
235Exaone 4.0 1.2B0.1%estimated ± 5.4 pp, low confidence
236Granite-4.0-H-1B0.1%estimated ± 5.4 pp, low confidence
237Gemma 3 27B0.1%estimated ± 5.4 pp, low confidence
238Granite-4.0-350M0.1%estimated ± 5.4 pp, low confidence
239Granite-4.0-H-350M0.1%estimated ± 5.4 pp, low confidence
240LFM2.5-VL-1.6B-Extract0.1%estimated ± 5.4 pp, low confidence
241Claude 3.5 Sonnet0.0%estimated ± 6.9 pp, low confidence
242LFM2.5-VL-450M0.0%estimated ± 6.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General