xAI · model
Grok Build 0.1 benchmark scores
As of 2026-10-07, Grok Build 0.1 (xAI) has measured scores on 1 benchmark and estimated scores on 14 more.
| Benchmark | Score | Source |
|---|---|---|
| GDPval-AA | 22.3% | estimated ± 8.8 pp, medium confidence |
| BrowseComp | 69.6% | estimated ± 5.7 pp, medium confidence |
| APEX-Agents-AA | 18.5% | estimated ± 11.4 pp, low confidence |
| CyberGym | 44.2% | estimated ± 10.0 pp, low confidence |
| τ³-bench results | 68.9% | estimated ± 8.9 pp, medium confidence |
| WideResearch | 72.8% | estimated ± 5.9 pp, low confidence |
| AA Agentic Index | 15.4% | estimated ± 6.4 pp, medium confidence |
| Gert Labs | 49.2% | measured |
| ApprenticeBench | 0.0% | estimated ± 3.1 pp, low confidence |
| OSWorld-Verified | 59.7% | estimated ± 5.9 pp, medium confidence |
| OSWorld 2.0 | 0.0% | estimated ± 4.5 pp, low confidence |
| Claw-Eval | 59.4% | estimated ± 6.8 pp, low confidence |
| MCP Atlas | 43.5% | estimated ± 14.8 pp, low confidence |
| Toolathlon | 41.2% | estimated ± 7.2 pp, medium confidence |
| ResearchClawBench | 15.7% | estimated ± 2.2 pp, medium confidence |