benchgap
Anthropic · model

Claude Sonnet 5 benchmark scores

As of 2026-10-07, Claude Sonnet 5 (Anthropic) has measured scores on 34 benchmarks and estimated scores on 83 more.

BenchmarkScoreSource
AA-GPQA Diamond91.1%measured
AA-HLE41.3%measured
AA-SciCode54.3%measured
Artificial Analysis Intelligence Index38.2%measured
CritPt16.9%measured
GDPval-AA48.3%measured
GPQA Diamond90.1%estimated ± 1.7 pp, high confidence
HLE57.4%measured
AA-LCR82.0%measured
LiveCodeBench v690.0%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy40.1%measured
SciCode41.5%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.180.4%measured
BrowseComp84.7%measured
HLE w/ tools57.4%measured
LongBench v261.5%estimated ± 5.9 pp, low confidence
VITA-Bench45.6%estimated ± 9.9 pp, low confidence
τ²-bench results88.6%estimated ± 8.8 pp, low confidence
AA Coding Index71.6%measured
AA-MMMU-Pro77.3%measured
APEX-Agents40.0%estimated ± 10.0 pp, low confidence
APEX-Agents-AA38.0%estimated ± 3.5 pp, low confidence
BioMysteryBench (human-difficult)41.0%estimated ± 5.9 pp, low confidence
DeepSearchQA92.5%estimated ± 4.5 pp, high confidence
SWE-bench Verified85.2%measured
AutomationBench28.7%estimated ± 9.7 pp, low confidence
BFCL v477.3%estimated ± 2.6 pp, medium confidence
CyberGym80.1%estimated ± 2.4 pp, medium confidence
JobBench49.9%estimated ± 8.6 pp, medium confidence
SWE-bench Pro63.2%measured
WideResearch82.8%estimated ± 3.4 pp, low confidence
AA Agentic Index44.3%measured
MMLU-Pro85.7%estimated ± 2.5 pp, low confidence
GPQA91.3%estimated ± 1.6 pp, high confidence
Gert Labs69.7%estimated ± 3.0 pp, medium confidence
AA-AnalystAgent51.5%estimated ± 2.5 pp, medium confidence
AA EnterpriseOps-Gym46.1%estimated ± 5.9 pp, low confidence
AA Harvey LAB92.9%estimated ± 1.2 pp, medium confidence
ApprenticeBench16.0%measured
ARC-AGI-195.4%estimated ± 0.7 pp, low confidence
ARC-AGI-262.5%estimated ± 3.1 pp, low confidence
CursorBench 3.154.6%estimated ± 3.4 pp, medium confidence
CursorBench 3.261.5%measured
Terminal-Bench 3.014.6%measured
FrontierCode 1.1 Main42.7%measured
FrontierSWE v211.1%estimated ± 7.6 pp, low confidence
OfficeQA Pro50.4%estimated ± 7.2 pp, medium confidence
OSWorld-Verified81.2%measured
Vals GPQA Diamond88.9%measured
Vals LiveCodeBench82.4%measured
Vals MMLU-Pro87.5%measured
Vals SWE-bench79.6%measured
Terminal-Bench 2.1 (Vals AI)74.5%measured
VulcanBench v375.2%estimated ± 0.7 pp, low confidence
AA AutomationBench26.9%estimated ± 5.7 pp, low confidence
GDP.pdf22.9%estimated ± 4.4 pp, low confidence
AA ITBench49.6%estimated ± 3.4 pp, medium confidence
AA Tau3 Banking47.4%estimated ± 5.8 pp, medium confidence
AA Terminal-Bench 2.184.6%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.019.0%estimated ± 8.9 pp, medium confidence
CursorBench 4.034.1%measured
CWE-bench v154.7%estimated ± 6.9 pp, medium confidence
DeepSWE65.7%estimated ± 3.4 pp, low confidence
HLE w/o tools43.2%measured
OSWorld 2.016.1%estimated ± 3.6 pp, high confidence
PostTrainBench v1.125.6%estimated ± 2.7 pp, high confidence
ProgramBench78.3%estimated ± 6.0 pp, low confidence
SWE Multilingual78.3%measured
Toolathlon-Verified70.9%estimated ± 2.3 pp, medium confidence
Vibe Code Bench71.8%estimated ± 5.6 pp, low confidence
CharXiv88.3%measured
CharXiv w/o tools77.0%measured
ExploitGym14.6%estimated ± 6.1 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval79.5%estimated ± 7.9 pp, medium confidence
MathVision80.1%estimated ± 2.1 pp, low confidence
MCP Atlas79.5%estimated ± 7.3 pp, medium confidence
MMLU-ProX86.0%estimated ± 1.7 pp, medium confidence
MMLU-Redux96.7%estimated ± 1.1 pp, medium confidence
MMMU-Pro73.7%estimated ± 1.4 pp, low confidence
NL2Repo45.6%estimated ± 4.7 pp, high confidence
NOVA-6358.5%estimated ± 1.6 pp, low confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
ScreenSpot Pro84.2%estimated ± 7.1 pp, medium confidence
SuperGPQA81.2%estimated ± 7.5 pp, medium confidence
Terminal-Bench 2.076.4%estimated ± 5.9 pp, medium confidence
Toolathlon54.9%estimated ± 1.7 pp, medium confidence
VideoMMMU86.7%estimated ± 1.5 pp, high confidence
ERQA70.4%estimated ± 3.7 pp, high confidence
HealthBench Hard32.9%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro86.2%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)78.4%estimated ± 6.4 pp, low confidence
MedXpertQA (Text)60.7%estimated ± 8.9 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals80.6%estimated ± 2.4 pp, medium confidence
ResearchClawBench20.5%estimated ± 2.0 pp, medium confidence
SWE-Rebench61.8%estimated ± 3.6 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)88.4%estimated ± 0.6 pp, medium confidence
DRACO81.4%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended55.1%estimated ± 1.6 pp, medium confidence
HealthBench (raw)50.3%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, medium confidence
HealthBench Professional59.0%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)60.5%estimated ± 5.0 pp, medium confidence
HLE-Verified31.0%measured
LABBench280.1%measured
MMMU84.9%estimated ± 0.9 pp, high confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, high confidence
SimpleVQA71.0%estimated ± 7.8 pp, low confidence
MMLU93.4%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python80.1%estimated ± 1.6 pp, high confidence
MathVision w/ Python82.9%estimated ± 0.9 pp, low confidence
RealWorldQA86.9%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)89.1%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.592.0%estimated ± 6.4 pp, low confidence