benchgap
xAI · model

Grok 4.20 benchmark scores

As of 2026-10-07, Grok 4.20 (xAI) has measured scores on 23 benchmarks and estimated scores on 66 more.

BenchmarkScoreSource
AA-GPQA Diamond88.7%estimated ± 1.2 pp, medium confidence
AA-HLE40.7%estimated ± 2.6 pp, medium confidence
AA-SciCode46.1%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index26.9%estimated ± 5.1 pp, medium confidence
CritPt7.8%estimated ± 4.9 pp, high confidence
GDPval-AA15.6%estimated ± 8.8 pp, medium confidence
GPQA Diamond88.5%measured
HLE44.8%estimated ± 5.5 pp, low confidence
LiveCodeBench v685.6%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy47.6%estimated ± 4.6 pp, medium confidence
SciCode45.3%estimated ± 2.9 pp, medium confidence
Terminal-Bench 2.146.8%estimated ± 6.9 pp, low confidence
BrowseComp58.7%estimated ± 5.7 pp, medium confidence
AA Coding Index53.8%estimated ± 5.1 pp, medium confidence
AA-MMMU-Pro73.2%estimated ± 2.2 pp, high confidence
APEX-Agents-AA10.5%estimated ± 11.4 pp, low confidence
DeepSearchQA62.8%measured
SWE-bench Verified76.7%measured
CyberGym30.1%estimated ± 10.0 pp, low confidence
SWE-bench Pro51.8%measured
τ³-bench results97.5%estimated ± 8.9 pp, low confidence
WideResearch54.0%estimated ± 5.9 pp, low confidence
AA Agentic Index10.2%estimated ± 6.4 pp, medium confidence
MMLU-Pro85.3%estimated ± 2.5 pp, medium confidence
GPQA88.5%estimated ± 0.1 pp, high confidence
Gert Labs38.4%measured
ApprenticeBench0.0%estimated ± 3.1 pp, low confidence
ARC-AGI-186.7%estimated ± 1.9 pp, high confidence
ARC-AGI-253.3%measured
CursorBench 3.149.0%estimated ± 5.6 pp, low confidence
CursorBench 3.244.1%estimated ± 3.6 pp, medium confidence
Terminal-Bench 3.00.9%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main14.9%estimated ± 2.7 pp, low confidence
FrontierSWE v20.5%estimated ± 9.0 pp, low confidence
OSWorld-Verified53.8%estimated ± 5.9 pp, medium confidence
Vals GPQA Diamond88.6%measured
Vals LiveCodeBench84.3%measured
Vals MMLU-Pro86.3%measured
Vals SWE-bench72.2%measured
Terminal-Bench 2.1 (Vals AI)44.2%measured
VulcanBench v383.2%estimated ± 5.6 pp, low confidence
AA Terminal-Bench 2.142.1%estimated ± 4.4 pp, medium confidence
AA Terminal-Bench 4.00.0%estimated ± 14.0 pp, low confidence
CursorBench 4.029.2%estimated ± 3.6 pp, medium confidence
DeepSWE56.3%estimated ± 7.2 pp, low confidence
HLE w/o tools31.6%measured
OSWorld 2.00.0%estimated ± 4.5 pp, low confidence
PostTrainBench v1.120.4%estimated ± 2.7 pp, high confidence
Vibe Code Bench4.1%measured
CharXiv60.9%measured
CharXiv w/o tools79.2%estimated ± 1.8 pp, medium confidence
Claw-Eval54.5%estimated ± 6.8 pp, low confidence
MathVision86.6%estimated ± 1.4 pp, low confidence
MCP Atlas43.5%estimated ± 14.8 pp, low confidence
MMLU-Redux92.8%estimated ± 3.0 pp, medium confidence
MMMU-Pro75.2%measured
NL2Repo41.6%estimated ± 4.7 pp, high confidence
ScreenSpot Pro85.0%estimated ± 2.5 pp, low confidence
SuperGPQA73.9%estimated ± 13.9 pp, low confidence
Terminal-Bench 2.047.1%measured
Toolathlon33.7%estimated ± 7.2 pp, low confidence
VideoMMMU84.2%estimated ± 1.0 pp, high confidence
ERQA54.1%measured
HealthBench Hard20.3%measured
LiveCodeBench Pro74.2%measured
MedXpertQA (MM)65.8%measured
MedXpertQA (Text)50.2%measured
MMLU-Pro (Arcee)86.4%estimated ± 3.4 pp, medium confidence
React Native Evals68.9%estimated ± 2.4 pp, low confidence
ResearchClawBench14.0%estimated ± 2.2 pp, medium confidence
SWE-Rebench59.0%estimated ± 3.6 pp, medium confidence
ARC-AGI-30.1%measured
FrontierCode 1.1 Extended18.6%estimated ± 2.4 pp, low confidence
HealthBench (length-adjusted)53.8%estimated ± 0.9 pp, low confidence
HealthBench Professional56.6%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)62.0%estimated ± 5.3 pp, low confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench281.1%estimated ± 1.8 pp, low confidence
MMMU82.7%estimated ± 1.1 pp, medium confidence
LiveCodeBench84.4%estimated ± 10.2 pp, low confidence
OpenHarmony Bench51.8%estimated ± 3.2 pp, high confidence
SimpleVQA57.4%measured
ZeroBench46.2%estimated ± 6.4 pp, low confidence
MMMLU89.3%estimated ± 1.1 pp, medium confidence
MMMU-Pro w/ Python76.7%estimated ± 0.6 pp, high confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA77.3%estimated ± 1.1 pp, low confidence
Video-MME (with subtitle)83.5%estimated ± 0.9 pp, low confidence
OmniDocBench 1.589.5%estimated ± 5.9 pp, low confidence