xAI · model
Grok 4.1 benchmark scores
As of 2026-10-07, Grok 4.1 (xAI) has measured scores on 1 benchmark and estimated scores on 10 more.
| Benchmark | Score | Source |
|---|---|---|
| GDPval-AA | 26.9% | estimated ± 8.3 pp, low confidence |
| BrowseComp | 46.0% | estimated ± 7.1 pp, low confidence |
| APEX-Agents-AA | 26.2% | estimated ± 10.0 pp, low confidence |
| DeepSearchQA | 72.6% | estimated ± 10.2 pp, low confidence |
| AA Agentic Index | 15.5% | estimated ± 6.8 pp, medium confidence |
| Gert Labs | 48.3% | estimated ± 8.5 pp, low confidence |
| OSWorld 2.0 | 7.2% | estimated ± 5.1 pp, low confidence |
| Claw-Eval | 55.4% | estimated ± 3.2 pp, high confidence |
| MCP Atlas | 38.6% | estimated ± 13.5 pp, low confidence |
| Toolathlon | 17.3% | estimated ± 6.9 pp, low confidence |
| ResearchClawBench | 13.5% | measured |