xAI · model
Grok 3 Mini benchmark scores
As of 2026-10-10, Grok 3 Mini (xAI) has measured scores on 2 benchmarks and estimated scores on 23 more.
| Benchmark | Score | Source |
|---|---|---|
| AA-GPQA Diamond | 79.1% | estimated ± 4.7 pp, high confidence |
| AA-HLE | 15.5% | estimated ± 5.3 pp, medium confidence |
| Artificial Analysis Intelligence Index | 14.5% | estimated ± 6.3 pp, medium confidence |
| CritPt | 0.7% | estimated ± 3.8 pp, high confidence |
| GPQA Diamond | 80.3% | estimated ± 0.4 pp, high confidence |
| HLE | 17.4% | estimated ± 6.4 pp, medium confidence |
| AA-Omniscience Accuracy | 19.7% | estimated ± 8.6 pp, medium confidence |
| MMLU-Pro | 83.7% | estimated ± 4.0 pp, high confidence |
| GPQA | 80.3% | measured |
| AIME 2025 | 84.3% | estimated ± 3.5 pp, medium confidence |
| ARC-AGI-1 | 49.1% | estimated ± 6.9 pp, low confidence |
| ARC-AGI-2 | 0.0% | estimated ± 11.6 pp, low confidence |
| Vals GPQA Diamond | 70.9% | estimated ± 4.8 pp, medium confidence |
| Vals MMLU-Pro | 70.4% | estimated ± 1.2 pp, medium confidence |
| HLE w/o tools | 11.7% | estimated ± 4.9 pp, high confidence |
| C-Eval | 1.7% | estimated ± 1.4 pp, low confidence |
| MMLU-Redux | 92.4% | estimated ± 1.9 pp, high confidence |
| SuperGPQA | 62.0% | estimated ± 8.2 pp, low confidence |
| HealthBench Hard | 0.0% | estimated ± 5.1 pp, low confidence |
| ARC-AGI-3 | 0.0% | estimated ± 2.1 pp, medium confidence |
| HealthBench Professional | 31.9% | estimated ± 4.8 pp, low confidence |
| AIME 2024 | 89.5% | measured |
| MMLU | 88.5% | estimated ± 4.3 pp, medium confidence |
| SimpleQA | 49.4% | estimated ± 10.9 pp, low confidence |
| BBH | 78.7% | estimated ± 10.9 pp, low confidence |