benchgap
xAI · model

Grok 4.1 benchmark scores

As of 2026-10-07, Grok 4.1 (xAI) has measured scores on 1 benchmark and estimated scores on 10 more.

BenchmarkScoreSource
GDPval-AA26.9%estimated ± 8.3 pp, low confidence
BrowseComp46.0%estimated ± 7.1 pp, low confidence
APEX-Agents-AA26.2%estimated ± 10.0 pp, low confidence
DeepSearchQA72.6%estimated ± 10.2 pp, low confidence
AA Agentic Index15.5%estimated ± 6.8 pp, medium confidence
Gert Labs48.3%estimated ± 8.5 pp, low confidence
OSWorld 2.07.2%estimated ± 5.1 pp, low confidence
Claw-Eval55.4%estimated ± 3.2 pp, high confidence
MCP Atlas38.6%estimated ± 13.5 pp, low confidence
Toolathlon17.3%estimated ± 6.9 pp, low confidence
ResearchClawBench13.5%measured