benchgap
Anthropic · model

Claude Sonnet 4.6 benchmark scores

As of 2026-10-07, Claude Sonnet 4.6 (Anthropic) has measured scores on 37 benchmarks and estimated scores on 82 more.

BenchmarkScoreSource
AA-GPQA Diamond79.9%measured
AA-HLE13.3%measured
AA-SciCode52.2%estimated ± 1.7 pp, medium confidence
AIME2695.5%estimated ± 0.6 pp, medium confidence
Artificial Analysis Intelligence Index24.7%measured
CritPt0.9%measured
GDPval-AA34.8%estimated ± 5.1 pp, medium confidence
GPQA Diamond89.9%estimated ± 0.1 pp, high confidence
HLE49.0%measured
IFBench56.0%estimated ± 6.9 pp, low confidence
AA-LCR68.3%measured
LiveCodeBench v686.2%estimated ± 8.1 pp, low confidence
AA-Omniscience Accuracy38.6%measured
SciCode40.9%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.174.6%estimated ± 6.9 pp, medium confidence
BrowseComp82.8%estimated ± 2.2 pp, high confidence
HLE w/ tools28.1%estimated ± 5.0 pp, low confidence
IFEval89.0%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench44.9%estimated ± 10.7 pp, low confidence
τ²-bench results79.5%measured
AA Coding Index49.5%estimated ± 3.3 pp, medium confidence
AA-MMMU-Pro70.6%measured
APEX-Agents-AA32.4%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)40.2%estimated ± 5.9 pp, low confidence
DeepSearchQA73.4%estimated ± 7.5 pp, low confidence
SWE-bench Verified79.6%measured
AutomationBench28.4%estimated ± 5.8 pp, low confidence
BFCL v477.9%estimated ± 8.6 pp, low confidence
CyberGym65.2%measured
JobBench36.9%measured
SWE-bench Pro57.1%estimated ± 2.1 pp, high confidence
τ³-bench results68.9%estimated ± 8.9 pp, medium confidence
WideResearch78.4%estimated ± 1.2 pp, medium confidence
AA Agentic Index22.4%estimated ± 5.1 pp, medium confidence
MMLU-Pro79.2%measured
FrontierMath v2 (Tier 4)8.3%measured
FrontierMath v2 (Tiers 1-3)32.4%measured
GPQA89.9%measured
AA-IFBench41.2%measured
Gert Labs62.9%measured
AA-AnalystAgent26.7%estimated ± 5.5 pp, medium confidence
AA EnterpriseOps-Gym22.0%estimated ± 3.6 pp, low confidence
AA Harvey LAB88.9%estimated ± 1.3 pp, medium confidence
ApprenticeBench2.0%measured
ARC-AGI-186.0%measured
ARC-AGI-258.3%measured
CursorBench 3.148.8%measured
CursorBench 3.239.1%estimated ± 1.6 pp, low confidence
Terminal-Bench 3.04.4%estimated ± 6.9 pp, low confidence
FrontierCode 1.1 Main24.3%measured
FrontierSWE v29.0%estimated ± 9.5 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified72.1%measured
Vals GPQA Diamond85.6%measured
Vals LiveCodeBench82.1%measured
Vals MMLU-Pro87.3%measured
Vals SWE-bench77.4%measured
Terminal-Bench 2.1 (Vals AI)57.3%measured
VulcanBench v381.9%estimated ± 5.6 pp, low confidence
AA AutomationBench49.9%estimated ± 5.0 pp, medium confidence
GDP.pdf21.5%estimated ± 4.4 pp, low confidence
AA ITBench58.9%estimated ± 3.6 pp, low confidence
AA Tau3 Banking24.1%estimated ± 14.6 pp, low confidence
AA Terminal-Bench 2.173.0%estimated ± 4.4 pp, high confidence
AA Terminal-Bench 4.010.6%estimated ± 14.0 pp, low confidence
CursorBench 4.029.5%estimated ± 4.4 pp, low confidence
DeepSWE13.7%estimated ± 3.4 pp, low confidence
HLE w/o tools33.9%estimated ± 4.9 pp, high confidence
OSWorld 2.08.3%measured
PostTrainBench v1.124.4%estimated ± 2.7 pp, high confidence
Toolathlon-Verified62.8%estimated ± 1.0 pp, low confidence
Vibe Code Bench51.5%measured
CharXiv77.4%measured
CharXiv w/o tools77.3%estimated ± 2.3 pp, medium confidence
ExploitGym2.0%estimated ± 6.9 pp, low confidence
C-Eval100.0%estimated ± 0.7 pp, low confidence
Claw-Eval67.8%measured
HMMT Nov 202593.2%estimated ± 2.6 pp, medium confidence
MathVision84.6%estimated ± 2.6 pp, high confidence
MCP Atlas62.2%estimated ± 8.2 pp, medium confidence
MCP-Tasks88.3%estimated ± 6.4 pp, low confidence
MMAnswerBench83.8%estimated ± 1.3 pp, medium confidence
MMLU-Redux89.8%estimated ± 0.7 pp, high confidence
MMMU-Pro77.1%estimated ± 3.2 pp, high confidence
NL2Repo44.3%estimated ± 4.7 pp, high confidence
ScreenSpot Pro67.0%estimated ± 7.1 pp, medium confidence
SuperGPQA95.0%measured
Terminal-Bench 2.059.1%measured
Toolathlon51.9%estimated ± 1.7 pp, low confidence
VideoMMMU84.6%estimated ± 1.5 pp, high confidence
ERQA64.5%estimated ± 3.7 pp, high confidence
HealthBench Hard1.7%estimated ± 5.1 pp, low confidence
MedXpertQA (MM)74.0%estimated ± 6.4 pp, low confidence
MMLU-Pro (Arcee)80.5%estimated ± 4.0 pp, medium confidence
React Native Evals80.6%measured
ResearchClawBench18.3%estimated ± 1.3 pp, high confidence
SWE-Rebench60.7%measured
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
FrontierMath (legacy)33.1%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
FrontierCode 1.1 Extended26.1%estimated ± 2.4 pp, low confidence
HealthBench (raw)19.5%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.5%estimated ± 2.5 pp, low confidence
HealthBench Professional58.8%estimated ± 1.3 pp, medium confidence
HealthBench Professional (raw)32.4%estimated ± 5.0 pp, low confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench263.7%estimated ± 1.3 pp, low confidence
MMMU79.6%estimated ± 0.9 pp, high confidence
LiveCodeBench86.2%estimated ± 10.2 pp, low confidence
Agents' Last Exam24.0%estimated ± 2.1 pp, medium confidence
OpenHarmony Bench53.5%estimated ± 3.2 pp, high confidence
SimpleVQA65.1%estimated ± 7.8 pp, low confidence
MMLU93.1%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python75.8%estimated ± 1.6 pp, high confidence
MathVision w/ Python76.8%estimated ± 0.9 pp, low confidence
RealWorldQA84.6%estimated ± 1.1 pp, low confidence
Video-MME (with subtitle)86.9%estimated ± 0.9 pp, low confidence
OmniDocBench 1.50.3%estimated ± 6.4 pp, low confidence