benchgap
Agentic · tools

GDP.pdf leaderboard

As of 2026-10-07, the highest measured score on GDP.pdf is 31.0% by GPT-6.1 Sol. 133 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6.1 Sol31.0%measured
2GPT-6 Astra31.0%measured
3Claude Opus 527.1%estimated ± 3.8 pp, high confidence
4Muse Spark 1.326.6%measured
5Claude Fable 5.126.2%measured
6Claude Opus 5.526.2%measured
7Claude Sonnet 5.525.8%measured
8GPT-5.525.5%estimated ± 3.8 pp, high confidence
9GPT-6 Sol25.2%measured
10Qwen3.8 Max24.9%estimated ± 3.4 pp, high confidence
11Qwen3.8-Flash-Next24.4%estimated ± 3.4 pp, high confidence
12GPT-5.6 Sol24.4%estimated ± 3.8 pp, high confidence
13Claude Sonnet 522.9%estimated ± 4.4 pp, low confidence
14GPT-5.6 Terra22.9%estimated ± 4.4 pp, low confidence
15GPT-6 Luna22.8%measured
16Ling 3.1 Flash22.5%estimated ± 5.4 pp, medium confidence
17Qwen3.8 Max Preview22.5%estimated ± 5.4 pp, medium confidence
18Muse Spark 1.222.5%estimated ± 5.4 pp, medium confidence
19GPT-5.422.4%estimated ± 4.4 pp, low confidence
20Claude Opus 4.722.0%estimated ± 4.4 pp, low confidence
21GPT-5.6 Luna22.0%estimated ± 4.4 pp, low confidence
22Kimi K322.0%measured
23Gemini 4 Argon21.8%measured
24Claude Opus 4.621.8%estimated ± 4.4 pp, low confidence
25Claude Sonnet 4.621.5%estimated ± 4.4 pp, low confidence
26Gemini 3.8 Flash21.0%measured
27Claude Haiku 5.520.8%measured
28Claude Opus 4.820.6%estimated ± 4.9 pp, high confidence
29Grok 4.720.0%measured
30Claude Opus 4.7 (Adaptive)19.8%estimated ± 4.9 pp, high confidence
31Claude Mythos Preview19.3%estimated ± 5.9 pp, low confidence
32MiMo-V2.6-Pro19.2%measured
33Mistral Large 418.6%measured
34GPT-5.5 Pro18.5%estimated ± 6.6 pp, low confidence
35Muse Spark 1.118.1%estimated ± 4.9 pp, high confidence
36Atria Dawn Preview17.8%estimated ± 5.2 pp, low confidence
37Claude 4.1 Opus17.8%estimated ± 5.2 pp, low confidence
38Claude 4 Sonnet17.8%estimated ± 5.2 pp, low confidence
39Claude Haiku 4.517.8%estimated ± 5.2 pp, low confidence
40Claude Opus 4.517.8%estimated ± 5.2 pp, low confidence
41Claude Sonnet 4.517.8%estimated ± 5.2 pp, low confidence
42Gemini 3 Flash17.8%estimated ± 5.2 pp, low confidence
43Gemini 3 Pro17.8%estimated ± 5.2 pp, low confidence
44GPT-5.1-Codex17.8%estimated ± 5.2 pp, low confidence
45GPT-5.217.8%estimated ± 5.2 pp, low confidence
46GPT-5.2-Codex17.8%estimated ± 5.2 pp, low confidence
47GPT-5.3 Codex17.8%estimated ± 5.2 pp, low confidence
48GPT-5 (high)17.8%estimated ± 5.2 pp, low confidence
49Kimi K2.517.8%estimated ± 5.2 pp, low confidence
50Qwen3.5 Plus17.8%estimated ± 5.2 pp, low confidence
51Claude Fable 517.6%estimated ± 3.7 pp, low confidence
52Gemini 3.5 Flash17.3%estimated ± 3.7 pp, low confidence
53Grok 4.616.7%estimated ± 3.7 pp, low confidence
54Qwen3.8-27B16.6%measured
55GLM-5.3-Flash15.4%measured
56GPT-5.4 Pro15.3%estimated ± 6.6 pp, low confidence
57Step 5 Preview14.8%measured
58MiMo-V2.6-Flash13.9%estimated ± 3.4 pp, high confidence
59Gemini 3.7 Flash13.3%estimated ± 3.4 pp, high confidence
60DeepSeek V4 Pro 081313.0%estimated ± 3.4 pp, medium confidence
61DeepSeek V4.1 Flash12.8%measured
62Inkling12.8%measured
63DeepSeek V4 Flash 073112.8%estimated ± 3.4 pp, medium confidence
64Claude Mythos 512.2%estimated ± 6.6 pp, low confidence
65Hy4 preview11.6%estimated ± 3.4 pp, medium confidence
66GLM-5.311.2%measured
67Ornith-1.5-397B10.6%estimated ± 6.6 pp, low confidence
68Grok 4.510.2%estimated ± 5.4 pp, medium confidence
69Kimi K2.610.1%estimated ± 4.9 pp, high confidence
70Muse Glimmer 30B10.0%measured
71MiniMax M39.8%measured
72GLM-5.29.6%estimated ± 5.4 pp, medium confidence
73dots3-note Preview9.5%estimated ± 6.6 pp, low confidence
74A.X K29.4%estimated ± 5.4 pp, medium confidence
75Apodex 1.19.4%estimated ± 5.4 pp, medium confidence
76Apodex 1.1 Mini9.4%estimated ± 5.4 pp, medium confidence
77Celeris-19.4%estimated ± 5.4 pp, low confidence
78Command A+9.4%estimated ± 5.4 pp, low confidence
79DeepSeek V39.4%estimated ± 5.4 pp, low confidence
80DeepSeek V3 03249.4%estimated ± 5.4 pp, low confidence
81Gemini 2.5 Pro9.4%estimated ± 5.4 pp, low confidence
82Gemini 3.1 Pro9.4%estimated ± 5.4 pp, medium confidence
83Gemini 3.5 Flash-Lite9.4%estimated ± 5.4 pp, medium confidence
84Gemini 3.6 Flash9.4%estimated ± 5.4 pp, medium confidence
85Gemma 3 27B9.4%estimated ± 5.4 pp, low confidence
86Gemma 4 12B9.4%estimated ± 5.4 pp, low confidence
87Gemma 4 26B A4B9.4%estimated ± 5.4 pp, low confidence
88Gemma 4 31B9.4%estimated ± 5.4 pp, low confidence
89Gemma 4 E2B9.4%estimated ± 5.4 pp, low confidence
90Gemma 4 E4B9.4%estimated ± 5.4 pp, low confidence
91GLM-4.79.4%estimated ± 5.4 pp, medium confidence
92GLM-5.19.4%estimated ± 5.4 pp, medium confidence
93GPT-4.1 mini9.4%estimated ± 5.4 pp, low confidence
94GPT-4.1 nano9.4%estimated ± 5.4 pp, low confidence
95GPT-4o9.4%estimated ± 5.4 pp, low confidence
96GPT-4o mini9.4%estimated ± 5.4 pp, low confidence
97GPT-5.19.4%estimated ± 5.4 pp, medium confidence
98GPT-5.4 mini9.4%estimated ± 5.4 pp, medium confidence
99GPT-5.4 nano9.4%estimated ± 5.4 pp, medium confidence
100GPT-OSS 120B9.4%estimated ± 5.4 pp, low confidence
101GPT-OSS 20B9.4%estimated ± 5.4 pp, low confidence
102Granite 4.2 30B9.4%estimated ± 5.4 pp, low confidence
103Granite 4.2 3B9.4%estimated ± 5.4 pp, low confidence
104Granite 4.2 8B9.4%estimated ± 5.4 pp, low confidence
105Grok 4.39.4%estimated ± 5.4 pp, medium confidence
106Hy39.4%estimated ± 5.4 pp, medium confidence
107Hy3 Preview9.4%estimated ± 5.4 pp, medium confidence
108Inkling-Small9.4%estimated ± 5.4 pp, medium confidence
109K-Exaone9.4%estimated ± 5.4 pp, low confidence
110Kimi K2.5 (Reasoning)9.4%estimated ± 5.4 pp, medium confidence
111Kimi K2.7 Code9.4%estimated ± 5.4 pp, medium confidence
112LFM2.5-2.6B9.4%estimated ± 5.4 pp, low confidence
113Ling 2.6 Flash9.4%estimated ± 5.4 pp, low confidence
114Ling 3.0 Flash9.4%estimated ± 5.4 pp, medium confidence
115Ling 3.0 Flash FP89.4%estimated ± 5.4 pp, medium confidence
116Ling 3.0 Flash VL9.4%estimated ± 5.4 pp, medium confidence
117Ling 3.0 Tiny9.4%estimated ± 5.4 pp, low confidence
118Llama 4 Maverick9.4%estimated ± 5.4 pp, low confidence
119Llama 4 Scout9.4%estimated ± 5.4 pp, low confidence
120Mercury 2.59.4%estimated ± 5.4 pp, low confidence
121MiMo-V2.5-Pro9.4%estimated ± 5.4 pp, medium confidence
122MiMo-V2-Flash9.4%estimated ± 5.4 pp, low confidence
123MiniCPM5-2B9.4%estimated ± 5.4 pp, low confidence
124MiniMax M2.79.4%estimated ± 5.4 pp, medium confidence
125Mistral Large 39.4%estimated ± 5.4 pp, low confidence
126Mistral Medium 3.5 128B9.4%estimated ± 5.4 pp, low confidence
127Mistral Small 49.4%estimated ± 5.4 pp, low confidence
128Mistral Small 4 (Reasoning)9.4%estimated ± 5.4 pp, low confidence
129Muse Spark9.4%estimated ± 5.4 pp, medium confidence
130Nemotron 3.5 Lightning 30B A3B NVFP49.4%estimated ± 5.4 pp, low confidence
131Nemotron 3 Nano 30B9.4%estimated ± 5.4 pp, low confidence
132Nemotron 3 Nano Omni 30B A3B9.4%estimated ± 5.4 pp, low confidence
133Nemotron 3 Super 100B9.4%estimated ± 5.4 pp, low confidence
134North Mini Code9.4%estimated ± 5.4 pp, low confidence
135Quasar 438B9.4%estimated ± 5.4 pp, medium confidence
136Qwen3.5-122B-A10B9.4%estimated ± 5.4 pp, medium confidence
137Qwen3.6-27B9.4%estimated ± 5.4 pp, medium confidence
138Qwen3.6-35B-A3B9.4%estimated ± 5.4 pp, medium confidence
139Qwen3.6 Plus9.4%estimated ± 5.4 pp, medium confidence
140Qwen3.7 Max9.4%estimated ± 5.4 pp, medium confidence
141Solar Pro 39.4%estimated ± 5.4 pp, low confidence
142Solar Pro 49.4%estimated ± 5.4 pp, medium confidence
143Step 3.7 Flash9.4%estimated ± 5.4 pp, medium confidence
144Trinity-Large-Preview9.4%estimated ± 5.4 pp, low confidence
145Trinity-Large-Thinking9.4%estimated ± 5.4 pp, low confidence
146Ultravox v0.6 Llama 3.3 70B9.4%estimated ± 5.4 pp, low confidence
147Beam9.3%estimated ± 6.6 pp, low confidence
148Agents-A19.3%estimated ± 6.6 pp, low confidence
149Agents-A1-4B9.3%estimated ± 6.6 pp, low confidence
150Ornith-1.5-35B-A3B9.3%estimated ± 6.6 pp, low confidence
151LongCat-Flash-Lite-Sparse9.3%estimated ± 6.6 pp, low confidence
152Ornith-1.5-9B9.3%estimated ± 6.6 pp, low confidence
153Qwen3.5-27B9.3%estimated ± 6.6 pp, low confidence
154Qwen3.5-35B-A3B9.3%estimated ± 6.6 pp, low confidence
155Qwen3.5 397B9.3%estimated ± 6.6 pp, low confidence
156Qwen3.7 Plus7.1%estimated ± 4.9 pp, medium confidence
157Nemotron 3 Ultra5.0%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General