benchgap
Anthropic · model

Claude Opus 5.5 benchmark scores

As of 2026-10-07, Claude Opus 5.5 (Anthropic) has measured scores on 40 benchmarks and estimated scores on 69 more.

BenchmarkScoreSource
AA-GPQA Diamond96.2%estimated ± 2.1 pp, low confidence
AA-HLE61.4%measured
AA-SciCode66.9%measured
Artificial Analysis Intelligence Index57.6%measured
CritPt31.7%measured
GDPval-AA68.3%measured
GPQA Diamond96.0%estimated ± 1.0 pp, low confidence
HLE67.6%estimated ± 1.9 pp, high confidence
AA-LCR84.7%measured
LiveCodeBench v6100.0%estimated ± 2.1 pp, medium confidence
AA-Omniscience Accuracy66.2%measured
SciCode66.1%estimated ± 3.2 pp, medium confidence
Terminal-Bench 2.189.8%estimated ± 3.9 pp, medium confidence
BrowseComp86.3%estimated ± 2.2 pp, high confidence
HLE w/ tools67.7%measured
LongBench v261.5%estimated ± 5.9 pp, low confidence
τ²-bench results99.9%estimated ± 4.8 pp, low confidence
AA Coding Index83.9%estimated ± 1.0 pp, medium confidence
AA-MMMU-Pro87.7%measured
APEX-Agents52.3%estimated ± 10.0 pp, low confidence
APEX-Agents-AA38.4%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)50.0%measured
DeepSearchQA90.5%estimated ± 8.4 pp, low confidence
SWE-bench Verified96.5%estimated ± 0.8 pp, low confidence
AutomationBench40.0%measured
BFCL v478.6%estimated ± 2.6 pp, medium confidence
CyberGym94.9%estimated ± 2.4 pp, low confidence
JobBench65.2%estimated ± 3.4 pp, low confidence
SWE-bench Pro89.9%measured
WideResearch87.9%estimated ± 3.4 pp, low confidence
AA Agentic Index61.3%estimated ± 2.8 pp, medium confidence
MMLU-Pro86.1%estimated ± 2.5 pp, low confidence
GPQA96.0%estimated ± 1.0 pp, low confidence
Gert Labs75.8%estimated ± 3.7 pp, low confidence
AA-AnalystAgent56.3%measured
AA EnterpriseOps-Gym55.6%estimated ± 1.8 pp, low confidence
AA Harvey LAB91.2%measured
ARC-AGI-197.5%measured
ARC-AGI-291.7%measured
CursorBench 3.185.7%estimated ± 3.4 pp, low confidence
CursorBench 3.270.5%estimated ± 1.6 pp, low confidence
FrontierCode 1.1 Main54.4%measured
FrontierSWE v262.3%measured
OfficeQA Pro67.7%measured
OSWorld-Verified100.0%estimated ± 3.0 pp, medium confidence
Vals GPQA Diamond98.2%estimated ± 1.5 pp, low confidence
Vals LiveCodeBench89.3%estimated ± 1.9 pp, medium confidence
Vals MMLU-Pro94.0%estimated ± 1.5 pp, low confidence
Vals SWE-bench96.9%estimated ± 1.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)94.5%estimated ± 4.6 pp, low confidence
VulcanBench v387.0%estimated ± 0.7 pp, low confidence
AA AutomationBench69.5%measured
GDP.pdf26.2%measured
AA ITBench38.2%measured
AA Tau3 Banking49.3%estimated ± 3.0 pp, low confidence
AA Terminal-Bench 2.186.8%estimated ± 2.6 pp, medium confidence
AA Terminal-Bench 4.059.6%measured
CursorBench 4.057.8%measured
CWE-bench v167.0%measured
DeepSWE74.2%measured
HLE w/o tools64.4%measured
OSWorld 2.048.7%measured
PostTrainBench v1.149.3%measured
ProgramBench91.2%measured
SWE Multilingual93.9%measured
Terminal-Bench 4.066.4%measured
Terminal-Bench-Science 0.158.7%measured
Toolathlon-Verified77.8%measured
Vibe Code Bench81.7%estimated ± 5.6 pp, low confidence
CharXiv86.0%estimated ± 7.5 pp, low confidence
CharXiv w/o tools87.2%estimated ± 2.4 pp, medium confidence
ExploitGym24.0%estimated ± 5.0 pp, low confidence
C-Eval97.1%estimated ± 1.0 pp, low confidence
MathVision89.7%estimated ± 4.9 pp, medium confidence
MCP Atlas87.5%estimated ± 2.7 pp, low confidence
MMLU-ProX91.7%estimated ± 1.7 pp, low confidence
MMLU-Redux98.1%estimated ± 1.1 pp, low confidence
MMMU-Pro82.5%estimated ± 4.4 pp, low confidence
NL2Repo81.4%estimated ± 6.4 pp, low confidence
NOVA-6364.5%estimated ± 1.6 pp, low confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
ScreenSpot Pro87.8%estimated ± 10.7 pp, low confidence
SuperGPQA90.7%estimated ± 7.5 pp, low confidence
Toolathlon55.9%estimated ± 6.3 pp, low confidence
ERQA68.9%estimated ± 3.9 pp, medium confidence
HealthBench Hard35.2%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro100.0%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)77.9%estimated ± 6.5 pp, low confidence
MedXpertQA (Text)78.7%estimated ± 8.9 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals83.8%estimated ± 4.7 pp, low confidence
ResearchClawBench20.6%estimated ± 2.2 pp, low confidence
ARC-AGI-373.3%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)89.3%measured
DRACO89.6%estimated ± 3.0 pp, low confidence
FrontierCode 1.1 Extended63.6%measured
HealthBench (raw)68.1%measured
HealthBench (length-adjusted)60.6%measured
HealthBench Professional65.6%measured
HealthBench Professional (raw)77.1%measured
HLE-Verified54.6%estimated ± 0.7 pp, medium confidence
LABBench284.1%estimated ± 1.3 pp, medium confidence
MMMU94.0%estimated ± 0.9 pp, medium confidence
Agents' Last Exam50.8%estimated ± 6.3 pp, medium confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, medium confidence
MMLU96.5%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python86.2%estimated ± 1.6 pp, medium confidence
RealWorldQA91.4%estimated ± 1.2 pp, low confidence
OmniDocBench 1.592.0%estimated ± 6.4 pp, low confidence