benchgap
Agentic · tools

GDPval-AA leaderboard

As of 2026-10-07, the highest measured score on GDPval-AA is 68.3% by Claude Opus 5.5. 116 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1BTL-3100.0%estimated ± 5.8 pp, low confidence
2Claude Opus 5.568.3%measured
3Claude Sonnet 5.567.0%measured
4GLM-5.3-Flash63.2%estimated ± 4.6 pp, high confidence
5Claude Fable 5.162.9%measured
6Claude Opus 561.2%measured
7Grok 4.760.8%measured
8MiMo-V2.6-Pro59.3%measured
9Muse Spark 1.359.2%measured
10Qwen3.8 Max Preview58.6%measured
11GLM-5.357.6%measured
12Qwen3.8-Flash-Next56.6%measured
13Gemini 4 Argon56.3%measured
14Grok 4.656.1%measured
15Ling 3.1 Flash56.1%measured
16Claude Fable 555.6%measured
17GPT-5.6 Sol55.6%measured
18MiMo-V2.6-Flash55.5%measured
19DeepSeek V4.1 Flash55.0%measured
20DeepSeek V4 Pro 081354.5%measured
21Step 5 Preview54.3%measured
22GPT-6.1 Sol53.8%measured
23Fugu Cyber53.8%estimated ± 5.1 pp, medium confidence
24Gemini 3.8 Flash Cyber53.2%estimated ± 5.1 pp, medium confidence
25GPT-5.5 Pro52.3%estimated ± 8.4 pp, medium confidence
26GPT-6 Astra52.1%measured
27Kimi K351.8%measured
28GPT-5.4 Pro51.1%estimated ± 8.4 pp, medium confidence
29Claude Mythos 551.1%estimated ± 5.1 pp, medium confidence
30Gemini 3.5 Flash Cyber50.5%estimated ± 5.1 pp, medium confidence
31GPT-6 Sol50.5%measured
32Claude Mythos Preview50.5%estimated ± 5.1 pp, medium confidence
33Muse Spark 1.248.9%measured
34Holo3-35B-A3B48.8%estimated ± 8.8 pp, medium confidence
35Claude Sonnet 548.3%measured
36GPT-5.6 Luna48.2%measured
37Claude Opus 4.847.8%measured
38GPT-5.6 Terra47.7%measured
39DeepSeek V4 Flash 073146.9%measured
40GPT-6 Luna46.9%measured
41Gemini 3.8 Flash46.8%measured
42Mistral Large 446.2%measured
43Qwen3.8-27B46.2%measured
44Gemini 3.7 Flash44.6%measured
45Grok 4.544.5%measured
46Holo3-122B-A10B43.9%estimated ± 8.8 pp, medium confidence
47GLM-5.243.7%measured
48Claude Opus 4.743.3%estimated ± 7.5 pp, medium confidence
49Claude Opus 4.7 (Adaptive)42.8%measured
50GPT-5.542.7%measured
51Gemini 3.5 Flash42.2%measured
52Claude Haiku 5.541.8%estimated ± 6.9 pp, medium confidence
53Ling 3.0 Flash Fin39.5%estimated ± 8.0 pp, low confidence
54Gemini 3.6 Flash39.3%measured
55Claude Opus 4.6 (Adaptive)38.7%estimated ± 11.2 pp, low confidence
56UI-Mate-27B38.1%estimated ± 8.8 pp, medium confidence
57GPT-5.437.4%measured
58MiniMax M337.3%measured
59GPT-5.237.3%estimated ± 6.2 pp, medium confidence
60GPT-5.3 Codex36.8%estimated ± 6.2 pp, medium confidence
61Claude Opus 4.635.8%estimated ± 5.1 pp, medium confidence
62Hy3 Preview35.8%measured
63Muse Spark 1.135.7%measured
64Beam35.5%estimated ± 8.4 pp, medium confidence
65Agents-A135.4%estimated ± 8.1 pp, medium confidence
66Apodex 1.134.8%measured
67Apodex 1.1 Mini34.8%measured
68Claude Sonnet 4.634.8%estimated ± 5.1 pp, medium confidence
69Quasar 438B34.2%measured
70Ling 3.0 Flash VL33.2%measured
71K-EXAONE 2.033.1%estimated ± 9.4 pp, low confidence
72Nemotron 3 Ultra33.1%measured
73Claude Sonnet 4.532.8%estimated ± 6.2 pp, medium confidence
74Ornith-1.0-397B32.7%estimated ± 9.4 pp, low confidence
75MiMo-V2.532.6%estimated ± 8.3 pp, low confidence
76GPT-5.1-Codex31.8%estimated ± 6.2 pp, medium confidence
77GPT-5.2-Codex31.7%estimated ± 6.2 pp, medium confidence
78Qwen3.7 Max31.6%measured
79Inkling-Small31.2%measured
80MiMo-V2.5-Pro31.2%measured
81GLM-5.131.0%measured
82Solar Pro 430.5%measured
83Qwen 3.6 Max (preview)30.1%estimated ± 12.5 pp, low confidence
84MiMo-V2-Pro29.9%estimated ± 8.3 pp, low confidence
85Grok 4.329.2%measured
86Claude 4.1 Opus29.0%estimated ± 6.2 pp, medium confidence
87Inkling28.9%measured
88Hy328.3%measured
89Ornith-1.0-35B27.9%estimated ± 9.4 pp, low confidence
90Solar Open 227.0%estimated ± 8.0 pp, low confidence
91Kimi K2.627.0%measured
92Kimi K2.7 Code27.0%measured
93Grok 4.126.9%estimated ± 8.3 pp, low confidence
94Grok 4.1 Fast (Reasoning)26.8%estimated ± 12.5 pp, low confidence
95Qwen3.5 Plus26.7%estimated ± 6.2 pp, medium confidence
96Claude 4 Sonnet26.6%estimated ± 6.2 pp, medium confidence
97Agents-A1-4B26.3%estimated ± 8.4 pp, medium confidence
98GLM-4.725.8%measured
99GPT-5.4 mini25.8%measured
100Step 3.7 Flash25.8%measured
101Laguna S 2.125.7%estimated ± 4.6 pp, medium confidence
102MiniMax M2.725.7%measured
103Hy4 preview25.4%estimated ± 3.7 pp, low confidence
104Muse Spark25.1%measured
105Claude Haiku 4.525.0%estimated ± 6.2 pp, medium confidence
106Atria Dawn Preview24.8%estimated ± 3.7 pp, low confidence
107Qwen3.8 Max24.8%estimated ± 3.7 pp, low confidence
108Qwen3.6 Plus24.7%measured
109Ornith-1.5-397B24.6%estimated ± 3.7 pp, low confidence
110Qwen3.6-27B24.4%measured
111Gemini 3.5 Flash-Lite24.3%measured
112dots3-note Preview24.0%estimated ± 3.7 pp, low confidence
113Ornith-1.0-9B24.0%estimated ± 9.4 pp, low confidence
114Claude Opus 4.523.4%estimated ± 3.7 pp, low confidence
115A.X K223.2%measured
116Qwen3.5-27B23.2%estimated ± 8.4 pp, medium confidence
117Qwen3.5-35B-A3B23.2%estimated ± 8.4 pp, medium confidence
118Qwen3.5 397B22.7%estimated ± 3.7 pp, low confidence
119GPT-5.4 nano22.5%measured
120Ling 3.0 Flash22.4%measured
121Ling 3.0 Flash FP822.4%measured
122Grok Build 0.122.3%estimated ± 8.8 pp, medium confidence
123Claude Opus 4.5 Thinking22.3%estimated ± 12.5 pp, low confidence
124Gemini 3 Flash21.9%estimated ± 6.2 pp, medium confidence
125Gemini 3 Pro21.9%estimated ± 6.2 pp, medium confidence
126GLM-521.6%estimated ± 3.7 pp, low confidence
127BTL-421.5%estimated ± 5.8 pp, medium confidence
128Grok 4.1 Fast21.0%estimated ± 8.8 pp, medium confidence
129Ornith-1.5-35B-A3B21.0%estimated ± 3.7 pp, low confidence
130GPT-5 (high)21.0%measured
131UI-Mate-9B20.9%estimated ± 8.8 pp, medium confidence
132GLM-5-Turbo20.2%estimated ± 9.4 pp, low confidence
133Qwen3.6-35B-A3B19.8%measured
134LongCat-Flash-Lite-Sparse19.8%estimated ± 8.4 pp, medium confidence
135GPT-5 (medium)19.0%estimated ± 12.5 pp, low confidence
136Qwen3 Max18.7%estimated ± 8.8 pp, medium confidence
137Ornith-1.5-9B18.7%estimated ± 3.7 pp, low confidence
138Grok 417.9%estimated ± 8.8 pp, medium confidence
139Kimi K2.517.2%measured
140Kimi K2.5 (Reasoning)17.2%measured
141GPT-5.116.5%measured
142Qwen3.5 397B (Reasoning)16.5%estimated ± 12.5 pp, low confidence
143Qwen3.5-122B-A10B15.8%measured
144GPT-5.1-Codex-Max15.6%estimated ± 12.5 pp, low confidence
145Gemini 3.1 Flash-Lite15.6%estimated ± 8.8 pp, medium confidence
146Grok 4.2015.6%estimated ± 8.8 pp, medium confidence
147MiMo-V2-Omni15.2%estimated ± 9.4 pp, low confidence
148Gemini 3.1 Pro14.7%measured
149Muse Glimmer 30B14.5%measured
150o313.6%estimated ± 12.5 pp, low confidence
151Qwen3.7 Plus13.5%measured
152Pokee-Isaac 28B13.3%estimated ± 5.8 pp, medium confidence
153Ternary Bonsai 2 27B13.2%estimated ± 12.5 pp, low confidence
154Mistral Medium 3.5 128B13.2%measured
155GLM-5V-Turbo11.6%estimated ± 8.8 pp, medium confidence
156DeepSeek V3.211.0%estimated ± 8.8 pp, low confidence
157GLM-4.610.7%estimated ± 12.5 pp, low confidence
158MiniCPM5-2B10.5%measured
159Grok Code Fast 19.9%estimated ± 12.5 pp, low confidence
160GPT-4.19.3%estimated ± 8.8 pp, low confidence
161Claude 4.1 Opus Thinking7.3%estimated ± 12.5 pp, low confidence
162Nemotron 3.5 Lightning 30B A3B NVFP47.0%measured
163MiMo-V2-Flash6.6%measured
164Gemma 4 31B6.1%measured
165GPT-OSS 120B5.6%measured
166Grok 4 Fast (Reasoning)4.8%estimated ± 12.5 pp, low confidence
167Granite 4.2 30B3.9%measured
168o13.6%estimated ± 12.5 pp, low confidence
169Gemma 4 26B A4B3.4%measured
170Ling 3.0 Tiny3.2%measured
171Kimi K23.2%estimated ± 12.5 pp, low confidence
172LLaDA2.2-flash3.0%estimated ± 5.8 pp, medium confidence
173Command A+0.9%measured
174Sarvam 105B0.7%estimated ± 12.5 pp, low confidence
175GLM-4.5-Air0.7%estimated ± 12.5 pp, low confidence
176DeepSeek V3.1 (Reasoning)0.2%estimated ± 12.5 pp, low confidence
177DeepSeek-R10.2%estimated ± 12.5 pp, low confidence
178DeepSeek V3.10.1%estimated ± 12.5 pp, low confidence
179Sarvam 30B0.1%estimated ± 12.5 pp, low confidence
180Solar Pro 20.1%estimated ± 12.5 pp, low confidence
181Mistral Large 20.1%estimated ± 12.5 pp, low confidence
182o3-mini0.0%estimated ± 12.5 pp, low confidence
183Mistral Medium 30.0%estimated ± 12.5 pp, low confidence
184Qwen3-Omni-30B-A3B-Thinking0.0%estimated ± 12.5 pp, low confidence
185Claude 3 Haiku0.0%estimated ± 12.5 pp, low confidence
186Exaone 4.0 1.2B0.0%estimated ± 12.5 pp, low confidence
187Granite-4.0-H-1B0.0%estimated ± 12.5 pp, low confidence
188Llama 3.1 405B0.0%estimated ± 12.5 pp, low confidence
189Qwen3-Omni-30B-A3B-Instruct0.0%estimated ± 12.5 pp, low confidence
190Gemini 2.5 Flash0.0%estimated ± 12.5 pp, low confidence
191Granite-4.0-H-350M0.0%estimated ± 12.5 pp, low confidence
192Nova Pro0.0%estimated ± 12.5 pp, low confidence
193Granite-4.0-350M0.0%estimated ± 12.5 pp, low confidence
194Nemotron Ultra 253B0.0%estimated ± 12.5 pp, low confidence
195LFM2.5-VL-1.6B-Extract0.0%estimated ± 12.5 pp, low confidence
196Celeris-10.0%measured
197DeepSeek V30.0%measured
198DeepSeek V3 03240.0%measured
199Exaone 4.0 32B0.0%estimated ± 12.5 pp, low confidence
200Gemini 2.5 Pro0.0%measured
201Gemma 3 27B0.0%measured
202Gemma 4 12B0.0%measured
203Gemma 4 E2B0.0%measured
204Gemma 4 E4B0.0%measured
205GPT-4.1 mini0.0%measured
206GPT-4.1 nano0.0%measured
207GPT-4o0.0%measured
208GPT-4o mini0.0%measured
209GPT-OSS 20B0.0%measured
210Granite 4.2 3B0.0%measured
211Granite 4.2 8B0.0%measured
212K-Exaone0.0%measured
213LFM2.5-2.6B0.0%measured
214LFM2.5-230M0.0%estimated ± 5.8 pp, low confidence
215LFM2.5-8B-A1B0.0%estimated ± 5.8 pp, low confidence
216LFM2.5-VL-3B0.0%estimated ± 5.8 pp, low confidence
217LFM2.5-VL-450M0.0%estimated ± 5.8 pp, low confidence
218Ling 2.6 Flash0.0%measured
219LLaDA2.2-mini0.0%estimated ± 5.8 pp, low confidence
220Llama 4 Maverick0.0%measured
221Llama 4 Scout0.0%measured
222Mellum2-12B-A2.5B-Instruct0.0%estimated ± 5.8 pp, low confidence
223Mellum2-12B-A2.5B-Thinking0.0%estimated ± 5.8 pp, low confidence
224Mercury 2.50.0%measured
225MiniCPM5-1B0.0%estimated ± 5.8 pp, low confidence
226Mistral Large 30.0%measured
227Mistral Small 40.0%measured
228Mistral Small 4 (Reasoning)0.0%measured
229Nemotron 3 Nano 30B0.0%measured
230Nemotron 3 Nano Omni 30B A3B0.0%measured
231Nemotron 3 Super 100B0.0%measured
232North Mini Code0.0%measured
233Phi-40.0%estimated ± 12.5 pp, low confidence
234Solar Pro 30.0%measured
235Trinity-Large-Preview0.0%measured
236Trinity-Large-Thinking0.0%measured
237Ultravox v0.6 Llama 3.3 70B0.0%measured
238ZAYA1-8B0.0%estimated ± 5.8 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General