benchgap
Vision & documents

GDP.pdf leaderboard

As of 2026-10-07, the highest measured score on GDP.pdf is 32.2% by GPT-6 Astra (xhigh). 3 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1GPT-6 Astra (xhigh)32.2%measured
2GPT-6.1 Sol (high)32.0%measured
3GPT-6.1 Sol (xhigh)31.8%measured
4GPT-6 Astra (high)31.0%measured
5GPT-6 Astra (max)31.0%measured
6GPT-6.1 Sol (max)31.0%measured
7GPT-6 Astra (medium)30.4%measured
8Claude Opus 5.5 (xhigh with fallback)29.1%estimated ± 4.0 pp, high confidence
9Claude Opus 5.5 (high with fallback)27.6%estimated ± 4.0 pp, high confidence
10Claude Opus 5.5 (medium with fallback)27.6%estimated ± 4.0 pp, high confidence
11Muse Spark 1.3 (max)26.6%measured
12Claude Fable 5.1 (max with fallback)26.2%measured
13Claude Opus 5.5 (max with fallback)26.2%measured
14Claude Sonnet 5.5 (max with fallback)25.8%measured
15GPT-6 Sol (max)25.2%measured
16Qwen3.8 Max (0902)22.8%measured
17GPT-6 Luna (max)22.8%measured
18Kimi K3 (max)22.0%measured
19Gemini 4 Argon (high)21.8%measured
20Gemini 3.8 Flash (high)21.0%measured
21Grok 4.7 (xhigh)20.0%measured
22MiMo-V2.6-Pro19.2%measured
23Mistral Large 4 Preview18.6%measured
24Qwen3.8 27B (xhigh)16.6%measured
25GLM-5.3-Flash15.4%measured
26Step 5 Preview14.8%measured
27Inkling (xhigh)12.8%measured
28DeepSeek V4.1 Flash (max)12.8%measured
29GLM-5.3 (max)11.2%measured
30Muse Glimmer (high)10.0%measured
31MiniMax-M39.8%measured
32K2 Horizon 375B A23B7.4%measured
33Nemotron 3 Ultra5.0%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents