benchgap
Anthropic · model

Claude 3.5 Sonnet benchmark scores

As of 2026-10-07, Claude 3.5 Sonnet (Anthropic) has measured scores on 4 benchmarks and estimated scores on 36 more.

BenchmarkScoreSource
AA-GPQA Diamond57.7%estimated ± 5.4 pp, medium confidence
AA-HLE4.9%estimated ± 6.5 pp, medium confidence
AA-SciCode32.6%estimated ± 3.9 pp, high confidence
AIME2695.5%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index10.4%estimated ± 6.2 pp, medium confidence
CritPt0.6%estimated ± 4.1 pp, high confidence
GPQA Diamond59.4%estimated ± 0.4 pp, high confidence
HLE7.8%estimated ± 8.2 pp, low confidence
LiveCodeBench v671.4%estimated ± 8.1 pp, low confidence
AA-Omniscience Accuracy16.3%estimated ± 8.1 pp, medium confidence
SciCode32.8%estimated ± 3.8 pp, high confidence
AA Coding Index23.7%estimated ± 5.1 pp, medium confidence
SWE-bench Verified49.0%measured
SWE-bench Pro23.8%estimated ± 3.2 pp, high confidence
MMLU-Pro71.8%estimated ± 3.0 pp, high confidence
FrontierMath v2 (Tier 4)0.0%measured
FrontierMath v2 (Tiers 1-3)2.1%measured
GPQA59.4%measured
ARC-AGI-10.0%estimated ± 6.9 pp, low confidence
ARC-AGI-20.0%estimated ± 11.6 pp, low confidence
FrontierCode 1.1 Main0.0%estimated ± 2.7 pp, low confidence
Vals GPQA Diamond42.3%estimated ± 4.8 pp, medium confidence
Vals LiveCodeBench19.8%estimated ± 9.9 pp, low confidence
Vals MMLU-Pro3.0%estimated ± 1.2 pp, medium confidence
Vals SWE-bench33.7%estimated ± 6.8 pp, low confidence
HLE w/o tools6.2%estimated ± 4.9 pp, medium confidence
Vibe Code Bench0.6%estimated ± 8.1 pp, low confidence
C-Eval73.0%estimated ± 0.9 pp, low confidence
HMMT Nov 202599.9%estimated ± 2.6 pp, low confidence
MMAnswerBench81.9%estimated ± 1.3 pp, low confidence
MMLU-Redux86.6%estimated ± 2.1 pp, high confidence
NL2Repo5.2%estimated ± 6.1 pp, low confidence
SuperGPQA36.4%estimated ± 9.2 pp, low confidence
HealthBench Hard0.0%estimated ± 5.1 pp, low confidence
SWE-Rebench57.0%estimated ± 3.6 pp, low confidence
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
FrontierMath (legacy)3.2%estimated ± 0.7 pp, low confidence
HealthBench Professional6.7%estimated ± 4.8 pp, low confidence
LiveCodeBench63.6%estimated ± 10.2 pp, low confidence
MMLU86.6%estimated ± 3.6 pp, medium confidence