benchgap
Anthropic · model

Claude Opus 4.8 benchmark scores

As of 2026-10-07, Claude Opus 4.8 (Anthropic) has measured scores on 47 benchmarks and estimated scores on 81 more.

BenchmarkScoreSource
AA-GPQA Diamond92.0%measured
AA-HLE48.7%measured
AA-SciCode54.4%measured
AIME26100.0%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index41.8%measured
CritPt20.9%measured
GDPval-AA47.8%measured
GPQA Diamond93.6%measured
HLE57.9%measured
IFBench62.6%estimated ± 6.9 pp, medium confidence
AA-LCR77.7%measured
LiveCodeBench v692.2%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy48.8%measured
SciCode52.3%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.174.6%measured
BrowseComp84.3%measured
HLE w/ tools57.5%estimated ± 4.6 pp, high confidence
IFEval92.9%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench49.0%estimated ± 9.9 pp, low confidence
τ²-bench results94.4%measured
AA Coding Index74.3%measured
AA-MMMU-Pro83.7%estimated ± 2.9 pp, high confidence
APEX-Agents35.6%estimated ± 0.6 pp, medium confidence
APEX-Agents-AA36.5%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)45.1%estimated ± 5.9 pp, low confidence
DeepSearchQA93.1%measured
SWE-bench Verified88.6%measured
AutomationBench28.6%estimated ± 9.7 pp, low confidence
BFCL v477.3%estimated ± 2.6 pp, medium confidence
CyberGym79.6%estimated ± 4.6 pp, high confidence
JobBench53.4%estimated ± 4.9 pp, medium confidence
SWE-bench Pro69.2%measured
τ³-bench results69.6%estimated ± 1.8 pp, low confidence
WideResearch79.7%estimated ± 4.6 pp, high confidence
AA Agentic Index42.6%measured
MMLU-Pro85.8%estimated ± 2.5 pp, low confidence
FrontierMath v2 (Tier 4)31.3%measured
FrontierMath v2 (Tiers 1-3)47.2%measured
GPQA93.6%measured
AA-IFBench62.2%measured
Gert Labs73.0%measured
AA-AnalystAgent54.1%estimated ± 5.5 pp, medium confidence
AA EnterpriseOps-Gym45.5%estimated ± 3.6 pp, medium confidence
AA Harvey LAB92.7%estimated ± 1.2 pp, medium confidence
ApprenticeBench20.0%estimated ± 3.1 pp, low confidence
ARC-AGI-192.5%measured
ARC-AGI-272.1%measured
CursorBench 3.158.4%measured
CursorBench 3.262.3%measured
Terminal-Bench 3.021.1%measured
FrontierCode 1.1 Main46.5%measured
FrontierSWE v231.5%estimated ± 9.0 pp, medium confidence
OfficeQA Pro66.2%measured
OSWorld-Verified83.4%measured
Vals GPQA Diamond92.4%measured
Vals LiveCodeBench87.8%measured
Vals MMLU-Pro89.6%measured
Vals SWE-bench88.6%measured
Terminal-Bench 2.1 (Vals AI)71.9%measured
VulcanBench v382.7%estimated ± 5.1 pp, low confidence
AA AutomationBench24.8%estimated ± 5.7 pp, low confidence
GDP.pdf20.6%estimated ± 4.9 pp, high confidence
AA ITBench49.9%estimated ± 3.6 pp, medium confidence
AA Tau3 Banking47.4%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.182.5%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.012.2%estimated ± 8.9 pp, medium confidence
CursorBench 4.038.2%estimated ± 1.6 pp, high confidence
CWE-bench v154.3%estimated ± 6.9 pp, medium confidence
DeepSWE69.4%estimated ± 3.4 pp, medium confidence
HLE w/o tools49.8%measured
OSWorld 2.020.6%measured
PostTrainBench v1.132.9%measured
ProgramBench87.9%estimated ± 4.8 pp, medium confidence
SWE Multilingual84.4%measured
Toolathlon-Verified73.4%estimated ± 1.0 pp, medium confidence
Vibe Code Bench75.2%estimated ± 5.6 pp, low confidence
CharXiv89.9%measured
CharXiv w/o tools80.5%measured
ExploitGym13.3%estimated ± 6.1 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval75.0%estimated ± 3.2 pp, medium confidence
DeepPlanning51.1%estimated ± 13.1 pp, low confidence
HMMT Nov 202593.2%estimated ± 2.6 pp, low confidence
MathVision86.0%estimated ± 2.1 pp, low confidence
MCP Atlas82.2%measured
MCP-Tasks77.1%estimated ± 0.6 pp, low confidence
MMAnswerBench89.0%estimated ± 1.3 pp, low confidence
MMLU-ProX88.3%estimated ± 1.7 pp, low confidence
MMLU-Redux97.3%estimated ± 1.1 pp, low confidence
MMMU-Pro74.4%estimated ± 1.4 pp, medium confidence
NL2Repo52.5%estimated ± 4.7 pp, high confidence
NOVA-6361.0%estimated ± 1.6 pp, low confidence
QwenClawBench67.7%estimated ± 2.6 pp, medium confidence
ScreenSpot Pro87.9%measured
SuperGPQA85.2%estimated ± 7.5 pp, low confidence
Terminal-Bench 2.074.8%estimated ± 5.9 pp, medium confidence
Toolathlon59.9%measured
VideoMMMU87.1%estimated ± 1.5 pp, high confidence
ERQA71.2%estimated ± 3.7 pp, high confidence
HealthBench Hard36.1%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro90.4%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)79.0%estimated ± 6.4 pp, low confidence
MedXpertQA (Text)66.3%estimated ± 4.5 pp, medium confidence
MMLU-Pro (Arcee)88.1%estimated ± 3.4 pp, low confidence
React Native Evals97.0%estimated ± 2.4 pp, low confidence
ResearchClawBench21.1%measured
SWE-Rebench61.8%estimated ± 3.6 pp, low confidence
ARC-AGI-31.5%measured
FrontierMath (legacy)47.7%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)89.2%estimated ± 0.6 pp, medium confidence
DRACO81.2%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended55.2%estimated ± 1.6 pp, medium confidence
HealthBench (raw)56.7%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.8%estimated ± 0.9 pp, medium confidence
HealthBench Professional59.9%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)66.6%estimated ± 5.0 pp, medium confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench275.9%estimated ± 1.3 pp, low confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, high confidence
SimpleVQA65.6%estimated ± 5.7 pp, low confidence
MMMLU90.9%estimated ± 1.1 pp, low confidence
MMLU94.8%estimated ± 1.1 pp, low confidence
MathVision w/ Python93.3%estimated ± 0.9 pp, low confidence
RealWorldQA87.3%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)89.4%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.590.8%estimated ± 7.4 pp, low confidence