benchgap
Anthropic · model

Claude Opus 4.6 (Adaptive) benchmark scores

As of 2026-10-07, Claude Opus 4.6 (Adaptive) (Anthropic) has measured scores on 13 benchmarks and estimated scores on 58 more.

BenchmarkScoreSource
AA-GPQA Diamond89.6%measured
AA-HLE39.9%measured
AA-SciCode53.5%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index32.0%measured
CritPt12.6%measured
GDPval-AA38.7%estimated ± 11.2 pp, low confidence
GPQA Diamond92.6%estimated ± 1.4 pp, high confidence
HLE48.5%estimated ± 6.3 pp, medium confidence
IFBench58.3%estimated ± 6.9 pp, medium confidence
AA-LCR78.0%measured
AA-Omniscience Accuracy47.0%measured
BrowseComp85.7%estimated ± 3.4 pp, high confidence
IFEval91.9%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
τ²-bench results92.1%measured
AA Coding Index67.8%estimated ± 6.8 pp, medium confidence
AA-MMMU-Pro75.4%measured
APEX-Agents-AA33.0%measured
BioMysteryBench (human-difficult)44.3%estimated ± 5.9 pp, low confidence
SWE-bench Verified81.7%estimated ± 4.4 pp, high confidence
BFCL v467.5%estimated ± 9.8 pp, low confidence
JobBench41.4%estimated ± 8.2 pp, low confidence
AA Agentic Index32.9%estimated ± 11.8 pp, low confidence
MMLU-Pro85.7%estimated ± 3.5 pp, high confidence
GPQA92.4%estimated ± 1.6 pp, high confidence
AA-IFBench53.1%measured
Gert Labs59.5%estimated ± 11.3 pp, low confidence
AA EnterpriseOps-Gym41.5%estimated ± 3.6 pp, medium confidence
ApprenticeBench9.2%estimated ± 5.0 pp, low confidence
ARC-AGI-193.0%measured
ARC-AGI-268.8%measured
FrontierCode 1.1 Main27.0%estimated ± 4.3 pp, medium confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
Vals GPQA Diamond92.1%estimated ± 2.2 pp, high confidence
Vals MMLU-Pro89.0%estimated ± 1.8 pp, high confidence
Vals SWE-bench80.2%estimated ± 6.3 pp, medium confidence
AA ITBench47.4%estimated ± 3.6 pp, medium confidence
HLE w/o tools38.3%estimated ± 5.7 pp, medium confidence
PostTrainBench v1.121.1%estimated ± 10.5 pp, low confidence
Vibe Code Bench53.5%measured
CharXiv82.3%estimated ± 7.5 pp, low confidence
CharXiv w/o tools79.8%estimated ± 2.4 pp, high confidence
C-Eval94.5%estimated ± 1.0 pp, low confidence
MathVision87.5%estimated ± 4.9 pp, high confidence
MCP Atlas72.3%estimated ± 9.4 pp, medium confidence
MCP-Tasks68.2%estimated ± 8.2 pp, low confidence
MMLU-Redux97.2%estimated ± 1.1 pp, low confidence
MMMU-Pro77.1%estimated ± 4.4 pp, medium confidence
ScreenSpot Pro74.1%estimated ± 10.7 pp, low confidence
SuperGPQA84.5%estimated ± 7.5 pp, low confidence
Toolathlon50.3%estimated ± 7.0 pp, low confidence
ERQA64.0%estimated ± 3.9 pp, high confidence
HealthBench Hard28.9%estimated ± 7.7 pp, low confidence
MedXpertQA (MM)72.8%estimated ± 6.5 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals82.2%estimated ± 3.0 pp, medium confidence
ResearchClawBench16.5%estimated ± 2.6 pp, medium confidence
SWE-Rebench62.1%estimated ± 6.5 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
BioMysteryBench (human-solvable)85.7%estimated ± 0.6 pp, low confidence
HealthBench (raw)49.0%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)56.3%estimated ± 2.5 pp, low confidence
HealthBench Professional55.1%estimated ± 4.6 pp, high confidence
HealthBench Professional (raw)59.3%estimated ± 5.0 pp, medium confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench276.8%estimated ± 1.3 pp, low confidence
MMMU83.3%estimated ± 0.9 pp, high confidence
MMLU94.5%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python78.9%estimated ± 1.6 pp, high confidence
RealWorldQA85.2%estimated ± 1.2 pp, medium confidence
OmniDocBench 1.590.0%estimated ± 6.4 pp, low confidence