benchgap
xAI · model

Grok 4.7 (xhigh) benchmark scores

As of 2026-10-07, Grok 4.7 (xhigh) (xAI) has measured scores on 7 benchmarks and estimated scores on 8 more.

BenchmarkScoreSource
Terminal-Bench 2.186.8%estimated ± 4.0 pp, high confidence
Terminal-Bench 4.025.8%measured
Terminal-Bench Science 0.13.8%estimated ± 6.5 pp, medium confidence
GPQA Diamond92.5%estimated ± 1.3 pp, high confidence
SciCode57.4%measured
CritPt17.7%measured
GDP.pdf20.0%measured
τ³-Bench Banking48.0%estimated ± 2.6 pp, medium confidence
MMMU-Pro80.0%estimated ± 2.6 pp, high confidence
AA-AnalystAgent52.7%estimated ± 7.8 pp, medium confidence
Harvey LAB92.9%estimated ± 1.6 pp, low confidence
Humanity's Last Exam43.1%measured
AutomationBench65.6%measured
EnterpriseOps-Gym42.3%estimated ± 6.2 pp, low confidence
ITBench SRE42.1%measured