benchgap
Anthropic · model

Claude Sonnet 4.5 Thinking benchmark scores

As of 2026-10-07, Claude Sonnet 4.5 Thinking (Anthropic) has measured scores on 3 benchmarks and estimated scores on 24 more.

BenchmarkScoreSource
AA-GPQA Diamond85.5%estimated ± 2.7 pp, high confidence
AA-HLE29.7%estimated ± 6.2 pp, medium confidence
AA-SciCode49.0%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index21.7%estimated ± 7.0 pp, medium confidence
CritPt9.8%estimated ± 5.7 pp, medium confidence
GPQA Diamond84.3%estimated ± 1.4 pp, medium confidence
HLE43.8%estimated ± 6.3 pp, medium confidence
AA-Omniscience Accuracy33.0%estimated ± 6.2 pp, medium confidence
AA Coding Index51.0%estimated ± 6.8 pp, medium confidence
SWE-bench Verified75.9%estimated ± 4.4 pp, high confidence
GPQA84.8%estimated ± 1.6 pp, high confidence
ARC-AGI-163.7%measured
ARC-AGI-213.6%measured
FrontierCode 1.1 Main9.4%estimated ± 4.3 pp, low confidence
Vals GPQA Diamond80.7%estimated ± 2.2 pp, high confidence
Vals MMLU-Pro82.3%estimated ± 1.8 pp, high confidence
Vals SWE-bench73.3%estimated ± 6.3 pp, medium confidence
HLE w/o tools28.2%estimated ± 6.0 pp, medium confidence
PostTrainBench v1.119.9%estimated ± 10.5 pp, low confidence
Vibe Code Bench22.6%measured
React Native Evals70.7%estimated ± 3.0 pp, low confidence
SWE-Rebench42.8%estimated ± 6.5 pp, low confidence
ARC-AGI-30.0%estimated ± 13.8 pp, low confidence
HealthBench (length-adjusted)43.6%estimated ± 2.5 pp, low confidence
HealthBench Professional55.1%estimated ± 4.6 pp, medium confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench222.7%estimated ± 1.3 pp, low confidence