benchgap
xAI · model

Grok 4.5 benchmark scores

As of 2026-10-07, Grok 4.5 (xAI) has measured scores on 27 benchmarks and estimated scores on 81 more.

BenchmarkScoreSource
AA-GPQA Diamond93.1%measured
AA-HLE42.7%measured
AA-SciCode55.0%measured
Artificial Analysis Intelligence Index38.8%measured
CritPt15.4%measured
GDPval-AA44.5%measured
GPQA Diamond89.5%estimated ± 1.4 pp, high confidence
HLE55.0%estimated ± 5.5 pp, low confidence
AA-LCR79.3%measured
LiveCodeBench v691.6%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy51.6%measured
SciCode51.5%estimated ± 2.9 pp, medium confidence
Terminal-Bench 2.183.3%measured
BrowseComp84.2%estimated ± 9.0 pp, medium confidence
HLE w/ tools54.4%estimated ± 6.1 pp, medium confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
τ²-bench results94.1%estimated ± 13.3 pp, low confidence
AA Coding Index72.5%measured
AA-MMMU-Pro80.4%measured
APEX-Agents37.5%estimated ± 10.0 pp, low confidence
APEX-Agents-AA33.5%estimated ± 9.1 pp, medium confidence
BioMysteryBench (human-difficult)46.3%estimated ± 5.9 pp, low confidence
DeepSearchQA92.1%estimated ± 4.5 pp, high confidence
SWE-bench Verified85.8%estimated ± 2.5 pp, high confidence
AutomationBench28.4%estimated ± 9.7 pp, low confidence
BFCL v477.0%estimated ± 2.6 pp, medium confidence
CyberGym77.9%estimated ± 5.3 pp, medium confidence
JobBench45.2%estimated ± 8.6 pp, medium confidence
SWE-bench Pro64.7%measured
WideResearch95.0%estimated ± 9.1 pp, low confidence
AA Agentic Index42.1%measured
MMLU-Pro85.7%estimated ± 3.5 pp, medium confidence
GPQA93.2%estimated ± 0.8 pp, high confidence
Gert Labs74.0%estimated ± 8.5 pp, medium confidence
AA-AnalystAgent34.9%estimated ± 12.0 pp, low confidence
AA EnterpriseOps-Gym42.7%estimated ± 6.7 pp, low confidence
AA Harvey LAB92.6%estimated ± 1.2 pp, medium confidence
ARC-AGI-185.7%measured
ARC-AGI-252.6%measured
CursorBench 3.156.3%estimated ± 3.4 pp, medium confidence
CursorBench 3.266.7%measured
Terminal-Bench 3.015.7%measured
FrontierCode 1.1 Main49.9%estimated ± 1.6 pp, medium confidence
FrontierSWE v223.5%estimated ± 9.0 pp, medium confidence
OfficeQA Pro59.6%estimated ± 7.2 pp, medium confidence
OSWorld-Verified78.9%estimated ± 8.7 pp, low confidence
Vals GPQA Diamond92.9%measured
Vals LiveCodeBench87.4%measured
Vals MMLU-Pro89.2%measured
Vals SWE-bench86.6%measured
Terminal-Bench 2.1 (Vals AI)67.8%measured
VulcanBench v389.9%measured
AA AutomationBench49.3%estimated ± 7.0 pp, medium confidence
GDP.pdf10.2%estimated ± 5.4 pp, medium confidence
AA Tau3 Banking42.5%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.184.8%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.022.6%estimated ± 8.9 pp, medium confidence
CursorBench 4.035.7%estimated ± 1.6 pp, high confidence
CWE-bench v152.0%estimated ± 6.9 pp, medium confidence
DeepSWE53.0%measured
HLE w/o tools44.5%estimated ± 5.0 pp, medium confidence
PostTrainBench v1.123.5%measured
ProgramBench0.0%estimated ± 4.8 pp, low confidence
SWE Multilingual78.0%measured
Toolathlon-Verified69.5%estimated ± 2.3 pp, medium confidence
Vibe Code Bench59.4%estimated ± 10.3 pp, low confidence
CharXiv85.8%estimated ± 7.5 pp, low confidence
CharXiv w/o tools82.9%estimated ± 2.4 pp, high confidence
ExploitGym18.2%estimated ± 12.2 pp, low confidence
C-Eval95.5%estimated ± 1.0 pp, low confidence
Claw-Eval80.7%estimated ± 13.6 pp, low confidence
MathVision89.7%estimated ± 4.9 pp, high confidence
MCP Atlas78.7%estimated ± 7.8 pp, low confidence
MMLU-ProX85.9%estimated ± 1.7 pp, medium confidence
MMLU-Redux97.5%estimated ± 1.1 pp, low confidence
MMMU-Pro79.4%estimated ± 4.4 pp, medium confidence
NL2Repo50.8%estimated ± 4.7 pp, high confidence
NOVA-6358.4%estimated ± 1.6 pp, low confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
ScreenSpot Pro79.6%estimated ± 10.7 pp, low confidence
SuperGPQA86.3%estimated ± 7.5 pp, low confidence
Terminal-Bench 2.072.1%estimated ± 5.9 pp, medium confidence
Toolathlon55.9%estimated ± 6.3 pp, medium confidence
ERQA68.5%estimated ± 3.9 pp, high confidence
HealthBench Hard33.1%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro87.3%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)77.1%estimated ± 6.5 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals93.1%estimated ± 2.4 pp, low confidence
ResearchClawBench20.3%estimated ± 2.0 pp, high confidence
SWE-Rebench58.1%estimated ± 8.9 pp, low confidence
ARC-AGI-30.3%measured
BioMysteryBench (human-solvable)87.8%estimated ± 0.6 pp, medium confidence
DRACO79.8%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended59.4%estimated ± 1.6 pp, medium confidence
HealthBench (raw)51.5%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.8%estimated ± 0.9 pp, medium confidence
HealthBench Professional59.9%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)61.7%estimated ± 5.0 pp, medium confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench263.1%estimated ± 1.3 pp, low confidence
MMMU87.6%estimated ± 0.9 pp, medium confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, high confidence
MMLU95.1%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python82.0%estimated ± 1.6 pp, high confidence
RealWorldQA87.8%estimated ± 1.2 pp, medium confidence
OmniDocBench 1.592.0%estimated ± 6.4 pp, low confidence