benchgap
Anthropic · model

Claude Opus 4.5 Thinking benchmark scores

As of 2026-10-07, Claude Opus 4.5 Thinking (Anthropic) has measured scores on 12 benchmarks and estimated scores on 52 more.

BenchmarkScoreSource
AA-GPQA Diamond86.6%measured
AA-HLE30.1%measured
AA-SciCode48.7%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index29.1%measured
CritPt4.6%measured
GDPval-AA22.3%estimated ± 12.5 pp, low confidence
GPQA Diamond87.6%estimated ± 1.4 pp, medium confidence
HLE44.0%estimated ± 6.3 pp, medium confidence
IFBench60.3%estimated ± 6.9 pp, medium confidence
AA-LCR77.3%measured
AA-Omniscience Accuracy46.6%measured
IFEval92.5%estimated ± 1.3 pp, high confidence
LongBench v261.5%estimated ± 5.9 pp, medium confidence
τ²-bench results89.5%measured
AA Coding Index49.9%estimated ± 6.8 pp, medium confidence
AA-MMMU-Pro74.0%measured
APEX-Agents-AA25.3%estimated ± 11.4 pp, low confidence
BioMysteryBench (human-difficult)44.1%estimated ± 5.9 pp, low confidence
SWE-bench Verified75.5%estimated ± 4.4 pp, high confidence
BFCL v464.9%estimated ± 9.8 pp, low confidence
AA Agentic Index22.5%estimated ± 12.3 pp, low confidence
MMLU-Pro85.7%estimated ± 3.5 pp, high confidence
GPQA88.9%estimated ± 1.6 pp, high confidence
AA-IFBench58.0%measured
AA EnterpriseOps-Gym37.0%estimated ± 3.6 pp, medium confidence
ARC-AGI-180.0%measured
ARC-AGI-237.6%measured
FrontierCode 1.1 Main8.5%estimated ± 4.3 pp, low confidence
OfficeQA Pro50.3%estimated ± 7.2 pp, low confidence
Vals GPQA Diamond87.3%estimated ± 2.2 pp, high confidence
Vals MMLU-Pro86.5%estimated ± 1.8 pp, high confidence
Vals SWE-bench72.6%estimated ± 6.3 pp, medium confidence
AA ITBench49.3%estimated ± 3.6 pp, medium confidence
HLE w/o tools27.2%estimated ± 5.7 pp, medium confidence
PostTrainBench v1.119.9%estimated ± 10.5 pp, low confidence
Vibe Code Bench20.6%measured
CharXiv79.2%estimated ± 7.5 pp, low confidence
CharXiv w/o tools78.9%estimated ± 2.4 pp, high confidence
C-Eval93.9%estimated ± 1.0 pp, low confidence
MathVision83.5%estimated ± 4.9 pp, high confidence
MCP-Tasks69.8%estimated ± 8.2 pp, low confidence
MMLU-Redux97.2%estimated ± 1.1 pp, low confidence
MMMU-Pro76.5%estimated ± 4.4 pp, medium confidence
ScreenSpot Pro72.7%estimated ± 10.7 pp, low confidence
SuperGPQA84.3%estimated ± 7.5 pp, low confidence
ERQA59.7%estimated ± 3.9 pp, high confidence
HealthBench Hard25.5%estimated ± 7.7 pp, low confidence
MedXpertQA (MM)69.7%estimated ± 6.5 pp, low confidence
MMLU-Pro (Arcee)94.4%estimated ± 4.0 pp, low confidence
React Native Evals69.0%estimated ± 3.0 pp, low confidence
SWE-Rebench34.7%estimated ± 6.5 pp, low confidence
ARC-AGI-31.7%estimated ± 3.4 pp, high confidence
BioMysteryBench (human-solvable)81.8%estimated ± 0.6 pp, low confidence
HealthBench (raw)39.3%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)51.1%estimated ± 2.5 pp, low confidence
HealthBench Professional55.1%estimated ± 4.6 pp, medium confidence
HealthBench Professional (raw)50.4%estimated ± 5.0 pp, low confidence
HLE-Verified35.3%estimated ± 0.7 pp, low confidence
LABBench251.6%estimated ± 1.3 pp, low confidence
MMMU82.2%estimated ± 0.9 pp, high confidence
MMLU94.5%estimated ± 1.1 pp, low confidence
MMMU-Pro w/ Python78.0%estimated ± 1.6 pp, high confidence
RealWorldQA84.4%estimated ± 1.2 pp, low confidence
OmniDocBench 1.567.2%estimated ± 6.4 pp, low confidence