benchgap
xAI · model

Grok 4.6 (high) benchmark scores

As of 2026-10-07, Grok 4.6 (high) (xAI) has measured scores on 4 benchmarks and estimated scores on 4 more.

BenchmarkScoreSource
GPQA Diamond94.9%measured
CritPt25.9%estimated ± 6.8 pp, medium confidence
τ³-Bench Banking50.7%measured
AA-AnalystAgent43.0%estimated ± 3.6 pp, medium confidence
Harvey LAB93.9%estimated ± 1.2 pp, medium confidence
Humanity's Last Exam50.4%estimated ± 4.6 pp, high confidence
AutomationBench66.7%measured
EnterpriseOps-Gym48.3%measured