benchgap
Anthropic · model

Claude Opus 4.6 benchmark scores

As of 2026-10-07, Claude Opus 4.6 (Anthropic) has measured scores on 41 benchmarks and estimated scores on 76 more.

BenchmarkScoreSource
AA-GPQA Diamond84.0%measured
AA-HLE19.1%measured
AA-SciCode46.9%estimated ± 3.8 pp, high confidence
AIME26100.0%estimated ± 0.6 pp, low confidence
Artificial Analysis Intelligence Index26.4%measured
CritPt2.8%measured
GDPval-AA35.8%estimated ± 5.1 pp, medium confidence
GPQA Diamond89.2%measured
HLE53.0%measured
IFBench56.4%estimated ± 6.9 pp, medium confidence
AA-LCR67.0%measured
LiveCodeBench v686.2%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy45.8%measured
SciCode48.9%estimated ± 3.8 pp, high confidence
BrowseComp83.7%measured
HLE w/ tools57.2%estimated ± 4.6 pp, high confidence
IFEval90.1%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
VITA-Bench45.4%estimated ± 9.9 pp, low confidence
τ²-bench results84.8%measured
AA Coding Index55.9%estimated ± 3.3 pp, medium confidence
AA-MMMU-Pro72.5%measured
APEX-Agents-AA33.8%estimated ± 3.5 pp, low confidence
BioMysteryBench (human-difficult)43.8%estimated ± 5.9 pp, low confidence
DeepSearchQA73.7%measured
SWE-bench Verified80.8%measured
AutomationBench28.4%estimated ± 5.8 pp, low confidence
BFCL v485.2%estimated ± 8.6 pp, low confidence
CyberGym66.6%measured
JobBench36.7%measured
SWE-bench Pro53.4%measured
τ³-bench results68.9%estimated ± 8.9 pp, medium confidence
WideResearch78.4%estimated ± 1.2 pp, medium confidence
AA Agentic Index23.7%estimated ± 5.1 pp, medium confidence
MMLU-Pro82.0%measured
FrontierMath v2 (Tier 4)22.9%measured
FrontierMath v2 (Tiers 1-3)40.7%measured
GPQA91.3%measured
AA-IFBench44.6%measured
Gert Labs61.9%measured
AA-AnalystAgent32.5%estimated ± 12.5 pp, low confidence
AA EnterpriseOps-Gym29.5%estimated ± 3.6 pp, medium confidence
AA Harvey LAB0.0%estimated ± 2.0 pp, low confidence
ApprenticeBench5.0%measured
ARC-AGI-187.9%estimated ± 2.7 pp, high confidence
ARC-AGI-252.6%estimated ± 10.1 pp, low confidence
CursorBench 3.150.1%estimated ± 5.6 pp, low confidence
CursorBench 3.242.4%estimated ± 1.6 pp, low confidence
FrontierCode 1.1 Main26.9%measured
FrontierSWE v23.4%estimated ± 9.0 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
OSWorld-Verified72.7%measured
Vals GPQA Diamond88.4%estimated ± 2.0 pp, high confidence
Vals LiveCodeBench81.8%estimated ± 1.9 pp, high confidence
Vals MMLU-Pro87.5%estimated ± 1.2 pp, high confidence
Vals SWE-bench81.6%estimated ± 1.4 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)57.4%estimated ± 6.3 pp, medium confidence
AA AutomationBench49.8%estimated ± 5.0 pp, medium confidence
GDP.pdf21.8%estimated ± 4.4 pp, low confidence
AA ITBench53.3%estimated ± 3.6 pp, low confidence
AA Tau3 Banking31.3%estimated ± 10.2 pp, low confidence
CursorBench 4.030.4%estimated ± 3.6 pp, medium confidence
DeepSWE21.8%estimated ± 3.4 pp, low confidence
HLE w/o tools40.0%measured
OSWorld 2.06.6%estimated ± 3.6 pp, high confidence
PostTrainBench v1.121.8%estimated ± 4.7 pp, medium confidence
Toolathlon-Verified73.2%estimated ± 1.3 pp, low confidence
Vibe Code Bench57.6%measured
CharXiv85.7%estimated ± 3.9 pp, high confidence
CharXiv w/o tools81.5%estimated ± 1.8 pp, medium confidence
ExploitGym11.4%estimated ± 6.1 pp, low confidence
C-Eval100.0%estimated ± 0.7 pp, low confidence
Claw-Eval70.4%measured
HMMT Nov 202593.2%estimated ± 2.6 pp, low confidence
MathVision86.2%estimated ± 1.4 pp, low confidence
MCP Atlas61.9%estimated ± 8.2 pp, medium confidence
MCP-Tasks93.7%estimated ± 6.4 pp, low confidence
MMAnswerBench87.1%estimated ± 1.3 pp, low confidence
MMLU-Redux91.4%estimated ± 0.7 pp, high confidence
MMMU-Pro77.3%measured
NL2Repo49.0%estimated ± 6.1 pp, medium confidence
ScreenSpot Pro83.1%measured
SuperGPQA95.0%measured
Terminal-Bench 2.065.4%measured
Toolathlon52.5%estimated ± 1.7 pp, low confidence
VideoMMMU84.5%estimated ± 1.0 pp, high confidence
ERQA51.6%measured
HealthBench Hard14.8%measured
LiveCodeBench Pro70.7%measured
MedXpertQA (MM)64.8%measured
MedXpertQA (Text)52.1%measured
MMLU-Pro (Arcee)89.1%measured
React Native Evals84.1%measured
ResearchClawBench19.9%measured
SWE-Rebench65.3%measured
ARC-AGI-30.0%estimated ± 2.1 pp, medium confidence
FrontierMath (legacy)41.3%estimated ± 0.7 pp, low confidence
BioMysteryBench (human-solvable)81.7%estimated ± 0.6 pp, low confidence
FrontierCode 1.1 Extended53.7%estimated ± 3.0 pp, low confidence
HealthBench (raw)26.9%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)53.9%estimated ± 3.7 pp, low confidence
HealthBench Professional52.7%estimated ± 4.8 pp, low confidence
HealthBench Professional (raw)39.0%estimated ± 5.0 pp, low confidence
HLE-Verified15.3%estimated ± 2.0 pp, low confidence
LABBench268.2%estimated ± 1.8 pp, low confidence
MMMU81.0%estimated ± 0.9 pp, high confidence
LiveCodeBench87.0%estimated ± 10.2 pp, low confidence
Agents' Last Exam24.0%estimated ± 2.1 pp, medium confidence
OpenHarmony Bench53.3%estimated ± 4.9 pp, low confidence
SimpleVQA72.8%estimated ± 5.7 pp, low confidence
ZeroBench66.4%estimated ± 9.3 pp, low confidence
MMMLU89.5%estimated ± 1.1 pp, medium confidence
MMLU94.4%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python78.7%estimated ± 0.6 pp, high confidence
RealWorldQA73.8%estimated ± 1.1 pp, low confidence
Video-MME (with subtitle)87.2%estimated ± 1.8 pp, medium confidence
OmniDocBench 1.591.7%estimated ± 5.9 pp, low confidence