benchgap
xAI · model

Grok 3 Mini benchmark scores

As of 2026-10-10, Grok 3 Mini (xAI) has measured scores on 2 benchmarks and estimated scores on 23 more.

BenchmarkScoreSource
AA-GPQA Diamond79.1%estimated ± 4.7 pp, high confidence
AA-HLE15.5%estimated ± 5.3 pp, medium confidence
Artificial Analysis Intelligence Index14.5%estimated ± 6.3 pp, medium confidence
CritPt0.7%estimated ± 3.8 pp, high confidence
GPQA Diamond80.3%estimated ± 0.4 pp, high confidence
HLE17.4%estimated ± 6.4 pp, medium confidence
AA-Omniscience Accuracy19.7%estimated ± 8.6 pp, medium confidence
MMLU-Pro83.7%estimated ± 4.0 pp, high confidence
GPQA80.3%measured
AIME 202584.3%estimated ± 3.5 pp, medium confidence
ARC-AGI-149.1%estimated ± 6.9 pp, low confidence
ARC-AGI-20.0%estimated ± 11.6 pp, low confidence
Vals GPQA Diamond70.9%estimated ± 4.8 pp, medium confidence
Vals MMLU-Pro70.4%estimated ± 1.2 pp, medium confidence
HLE w/o tools11.7%estimated ± 4.9 pp, high confidence
C-Eval1.7%estimated ± 1.4 pp, low confidence
MMLU-Redux92.4%estimated ± 1.9 pp, high confidence
SuperGPQA62.0%estimated ± 8.2 pp, low confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
HealthBench Professional31.9%estimated ± 4.8 pp, low confidence
AIME 202489.5%measured
MMLU88.5%estimated ± 4.3 pp, medium confidence
SimpleQA49.4%estimated ± 10.9 pp, low confidence
BBH78.7%estimated ± 10.9 pp, low confidence