Knowledge & reasoning
GPQA Diamond leaderboard
As of 2026-10-07, the highest measured score on GPQA Diamond is 96.3% by GPT-6 Astra (xhigh). 17 more models have estimated scores, calibrated from the benchmarks they were measured on.
Measured scores: artificial-analysis.
| # | Model | Score | Source |
|---|---|---|---|
| 1 | GPT-6 Astra (xhigh) | 96.3% | measured |
| 2 | GPT-6 Astra (max) | 96.1% | measured |
| 3 | Claude Opus 5.5 (max with fallback) | 96.0% | estimated ± 1.3 pp, medium confidence |
| 4 | Claude Fable 5.1 (xhigh with fallback) | 95.6% | estimated ± 1.3 pp, high confidence |
| 5 | Claude Opus 5.5 (xhigh with fallback) | 95.4% | estimated ± 1.3 pp, high confidence |
| 6 | Gemini 4 Argon (high) | 95.3% | estimated ± 1.3 pp, high confidence |
| 7 | Gemini 3.8 Flash (high) | 95.3% | measured |
| 8 | Claude Fable 5.1 (high with fallback) | 95.2% | estimated ± 1.3 pp, high confidence |
| 9 | Claude Opus 5.5 (high with fallback) | 95.1% | estimated ± 1.3 pp, high confidence |
| 10 | Claude Fable 5 (with fallback) | 95.1% | estimated ± 1.3 pp, high confidence |
| 11 | Claude Sonnet 5.5 (max with fallback) | 95.0% | estimated ± 1.3 pp, high confidence |
| 12 | GPT-6 Astra (high) | 94.9% | measured |
| 13 | Grok 4.6 (high) | 94.9% | measured |
| 14 | GPT-6.1 Sol (max) | 94.6% | estimated ± 1.3 pp, high confidence |
| 15 | Gemini 3.7 Flash (high) | 94.5% | measured |
| 16 | GPT-6.1 Sol (xhigh) | 94.2% | estimated ± 2.3 pp, high confidence |
| 17 | GPT-5.6 Sol (max) | 94.1% | measured |
| 18 | MiMo-V2.6-Pro | 94.0% | estimated ± 1.3 pp, high confidence |
| 19 | Claude Fable 5.1 (max with fallback) | 93.7% | measured |
| 20 | GPT-6 Sol (max) | 93.6% | estimated ± 1.3 pp, high confidence |
| 21 | Kimi K3 (max) | 93.5% | measured |
| 22 | Muse Spark 1.3 (max) | 93.5% | measured |
| 23 | Step 5 Preview | 93.3% | estimated ± 1.3 pp, high confidence |
| 24 | MiniMax-M3 | 92.9% | measured |
| 25 | Qwen3.8 Max (0902) | 92.8% | measured |
| 26 | Grok 4.7 (xhigh) | 92.5% | estimated ± 1.3 pp, high confidence |
| 27 | GLM-5.3 (max) | 91.7% | measured |
| 28 | DeepSeek V4.1 Flash (max) | 91.4% | estimated ± 1.3 pp, high confidence |
| 29 | GLM-5.3-Flash | 91.2% | measured |
| 30 | GPT-6 Luna (max) | 91.2% | estimated ± 1.3 pp, high confidence |
| 31 | Qwen3.8 27B (xhigh) | 90.5% | measured |
| 32 | Mistral Large 4 Preview | 90.0% | estimated ± 1.3 pp, high confidence |
| 33 | K2 Horizon 375B A23B | 87.3% | measured |
| 34 | Inkling (xhigh) | 87.2% | measured |
| 35 | Nemotron 3 Ultra | 86.7% | measured |
| 36 | Muse Glimmer (high) | 83.5% | measured |
All leaderboards
Agentic · terminal
Agentic · tools
- AA-AnalystAgent
- APEX-Agents
- AutomationBench
- EnterpriseOps-Gym
- Harvey LAB
- ITBench SRE
- τ³-Bench Banking
- τ²-Bench Telecom