benchgap
Knowledge & reasoning

GPQA Diamond leaderboard

As of 2026-10-07, the highest measured score on GPQA Diamond is 96.3% by GPT-6 Astra (xhigh). 17 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1GPT-6 Astra (xhigh)96.3%measured
2GPT-6 Astra (max)96.1%measured
3Claude Opus 5.5 (max with fallback)96.0%estimated ± 1.3 pp, medium confidence
4Claude Fable 5.1 (xhigh with fallback)95.6%estimated ± 1.3 pp, high confidence
5Claude Opus 5.5 (xhigh with fallback)95.4%estimated ± 1.3 pp, high confidence
6Gemini 4 Argon (high)95.3%estimated ± 1.3 pp, high confidence
7Gemini 3.8 Flash (high)95.3%measured
8Claude Fable 5.1 (high with fallback)95.2%estimated ± 1.3 pp, high confidence
9Claude Opus 5.5 (high with fallback)95.1%estimated ± 1.3 pp, high confidence
10Claude Fable 5 (with fallback)95.1%estimated ± 1.3 pp, high confidence
11Claude Sonnet 5.5 (max with fallback)95.0%estimated ± 1.3 pp, high confidence
12GPT-6 Astra (high)94.9%measured
13Grok 4.6 (high)94.9%measured
14GPT-6.1 Sol (max)94.6%estimated ± 1.3 pp, high confidence
15Gemini 3.7 Flash (high)94.5%measured
16GPT-6.1 Sol (xhigh)94.2%estimated ± 2.3 pp, high confidence
17GPT-5.6 Sol (max)94.1%measured
18MiMo-V2.6-Pro94.0%estimated ± 1.3 pp, high confidence
19Claude Fable 5.1 (max with fallback)93.7%measured
20GPT-6 Sol (max)93.6%estimated ± 1.3 pp, high confidence
21Kimi K3 (max)93.5%measured
22Muse Spark 1.3 (max)93.5%measured
23Step 5 Preview93.3%estimated ± 1.3 pp, high confidence
24MiniMax-M392.9%measured
25Qwen3.8 Max (0902)92.8%measured
26Grok 4.7 (xhigh)92.5%estimated ± 1.3 pp, high confidence
27GLM-5.3 (max)91.7%measured
28DeepSeek V4.1 Flash (max)91.4%estimated ± 1.3 pp, high confidence
29GLM-5.3-Flash91.2%measured
30GPT-6 Luna (max)91.2%estimated ± 1.3 pp, high confidence
31Qwen3.8 27B (xhigh)90.5%measured
32Mistral Large 4 Preview90.0%estimated ± 1.3 pp, high confidence
33K2 Horizon 375B A23B87.3%measured
34Inkling (xhigh)87.2%measured
35Nemotron 3 Ultra86.7%measured
36Muse Glimmer (high)83.5%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents