benchgap
Anthropic · model

Claude Sonnet 5.5 benchmark scores

As of 2026-10-07, Claude Sonnet 5.5 (Anthropic) has measured scores on 33 benchmarks and estimated scores on 65 more.

BenchmarkScoreSource
AA-GPQA Diamond96.2%estimated ± 2.1 pp, low confidence
AA-HLE55.0%measured
AA-SciCode61.0%measured
Artificial Analysis Intelligence Index56.0%measured
CritPt31.4%measured
GDPval-AA67.0%measured
GPQA Diamond96.0%estimated ± 1.0 pp, low confidence
HLE63.7%estimated ± 5.7 pp, medium confidence
AA-LCR82.7%measured
LiveCodeBench v699.0%estimated ± 2.1 pp, medium confidence
AA-Omniscience Accuracy54.0%measured
SciCode60.3%estimated ± 3.2 pp, medium confidence
Terminal-Bench 2.189.9%estimated ± 3.9 pp, medium confidence
BrowseComp87.0%estimated ± 6.4 pp, medium confidence
HLE w/ tools64.5%measured
LongBench v261.5%estimated ± 5.9 pp, low confidence
τ²-bench results94.3%estimated ± 7.1 pp, low confidence
AA Coding Index83.0%estimated ± 1.0 pp, medium confidence
AA-MMMU-Pro83.4%estimated ± 2.9 pp, high confidence
APEX-Agents51.5%estimated ± 10.0 pp, low confidence
APEX-Agents-AA48.1%estimated ± 9.1 pp, low confidence
BioMysteryBench (human-difficult)44.7%measured
DeepSearchQA90.5%estimated ± 8.4 pp, low confidence
SWE-bench Verified88.3%estimated ± 0.8 pp, medium confidence
AutomationBench40.4%estimated ± 7.6 pp, low confidence
BFCL v478.5%estimated ± 2.6 pp, medium confidence
CyberGym88.7%estimated ± 2.4 pp, low confidence
JobBench65.2%estimated ± 3.4 pp, low confidence
SWE-bench Pro81.3%measured
WideResearch86.4%estimated ± 3.4 pp, low confidence
AA Agentic Index60.3%estimated ± 2.8 pp, medium confidence
MMLU-Pro86.0%estimated ± 2.5 pp, low confidence
GPQA96.0%estimated ± 1.0 pp, low confidence
Gert Labs85.6%estimated ± 8.9 pp, low confidence
AA-AnalystAgent57.5%measured
AA EnterpriseOps-Gym56.3%estimated ± 1.8 pp, low confidence
AA Harvey LAB93.1%measured
ARC-AGI-197.2%estimated ± 3.7 pp, high confidence
ARC-AGI-284.8%estimated ± 13.0 pp, low confidence
CursorBench 3.171.1%estimated ± 3.4 pp, medium confidence
CursorBench 3.263.2%estimated ± 1.6 pp, medium confidence
FrontierCode 1.1 Main46.2%measured
FrontierSWE v261.9%measured
OfficeQA Pro65.6%measured
OSWorld-Verified90.7%estimated ± 8.7 pp, low confidence
Vals GPQA Diamond98.2%estimated ± 1.5 pp, low confidence
Vals LiveCodeBench87.0%estimated ± 1.9 pp, high confidence
Vals MMLU-Pro96.7%estimated ± 1.5 pp, low confidence
Vals SWE-bench95.4%estimated ± 1.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)98.6%estimated ± 4.6 pp, low confidence
VulcanBench v387.0%estimated ± 0.7 pp, low confidence
AA AutomationBench71.8%measured
GDP.pdf25.8%measured
AA ITBench43.4%estimated ± 3.4 pp, medium confidence
AA Tau3 Banking47.0%estimated ± 4.4 pp, medium confidence
AA Terminal-Bench 2.186.8%estimated ± 2.6 pp, medium confidence
AA Terminal-Bench 4.063.6%measured
CursorBench 4.055.5%measured
CWE-bench v165.0%estimated ± 6.0 pp, medium confidence
DeepSWE71.0%measured
HLE w/o tools56.9%measured
OSWorld 2.045.2%estimated ± 5.5 pp, low confidence
PostTrainBench v1.144.4%estimated ± 4.3 pp, medium confidence
ProgramBench79.7%measured
SWE Multilingual90.3%measured
Terminal-Bench 4.070.6%measured
Terminal-Bench-Science 0.159.9%measured
Toolathlon-Verified77.8%measured
Vibe Code Bench75.0%estimated ± 5.6 pp, low confidence
ExploitGym23.2%estimated ± 5.0 pp, low confidence
C-Eval97.0%estimated ± 1.0 pp, low confidence
MCP Atlas88.0%estimated ± 2.7 pp, low confidence
MMLU-ProX90.4%estimated ± 1.7 pp, low confidence
MMLU-Redux97.6%estimated ± 1.1 pp, low confidence
NL2Repo71.6%estimated ± 6.4 pp, low confidence
NOVA-6363.2%estimated ± 1.6 pp, low confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
SuperGPQA87.1%estimated ± 7.5 pp, low confidence
Toolathlon55.9%estimated ± 6.3 pp, low confidence
HealthBench Hard35.1%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro98.0%estimated ± 6.0 pp, low confidence
MedXpertQA (Text)72.9%estimated ± 8.9 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals81.3%estimated ± 4.7 pp, low confidence
ResearchClawBench20.5%estimated ± 2.2 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
BioMysteryBench (human-solvable)89.2%measured
DRACO87.0%measured
FrontierCode 1.1 Extended59.1%measured
HealthBench (raw)69.4%measured
HealthBench (length-adjusted)65.4%measured
HealthBench Professional69.2%measured
HealthBench Professional (raw)77.1%measured
HLE-Verified50.6%estimated ± 6.0 pp, low confidence
LABBench286.2%estimated ± 2.4 pp, medium confidence
Agents' Last Exam50.2%estimated ± 6.3 pp, medium confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, medium confidence
MMLU95.4%estimated ± 1.1 pp, low confidence