xAI · model
Grok 4.7 (xhigh) benchmark scores
As of 2026-10-07, Grok 4.7 (xhigh) (xAI) has measured scores on 7 benchmarks and estimated scores on 8 more.
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench 2.1 | 86.8% | estimated ± 4.0 pp, high confidence |
| Terminal-Bench 4.0 | 25.8% | measured |
| Terminal-Bench Science 0.1 | 3.8% | estimated ± 6.5 pp, medium confidence |
| GPQA Diamond | 92.5% | estimated ± 1.3 pp, high confidence |
| SciCode | 57.4% | measured |
| CritPt | 17.7% | measured |
| GDP.pdf | 20.0% | measured |
| τ³-Bench Banking | 48.0% | estimated ± 2.6 pp, medium confidence |
| MMMU-Pro | 80.0% | estimated ± 2.6 pp, high confidence |
| AA-AnalystAgent | 52.7% | estimated ± 7.8 pp, medium confidence |
| Harvey LAB | 92.9% | estimated ± 1.6 pp, low confidence |
| Humanity's Last Exam | 43.1% | measured |
| AutomationBench | 65.6% | measured |
| EnterpriseOps-Gym | 42.3% | estimated ± 6.2 pp, low confidence |
| ITBench SRE | 42.1% | measured |