benchgap
Anthropic · model

Claude Opus 5 (max) benchmark scores

As of 2026-10-07, Claude Opus 5 (max) (Anthropic) has measured scores on 2 benchmarks and estimated scores on 3 more.

BenchmarkScoreSource
τ³-Bench Banking47.1%estimated ± 4.5 pp, medium confidence
AA-AnalystAgent53.8%measured
Harvey LAB93.5%measured
AutomationBench67.5%estimated ± 11.2 pp, low confidence
EnterpriseOps-Gym53.6%estimated ± 1.8 pp, low confidence