Anthropic · model
Claude Opus 5.5 (max with fallback) benchmark scores
As of 2026-10-07, Claude Opus 5.5 (max with fallback) (Anthropic) has measured scores on 11 benchmarks and estimated scores on 4 more.
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench 2.1 | 88.2% | estimated ± 4.0 pp, medium confidence |
| Terminal-Bench 4.0 | 59.6% | measured |
| Terminal-Bench Science 0.1 | 59.0% | measured |
| GPQA Diamond | 96.0% | estimated ± 1.3 pp, medium confidence |
| SciCode | 66.9% | measured |
| CritPt | 31.7% | measured |
| GDP.pdf | 26.2% | measured |
| τ³-Bench Banking | 49.0% | estimated ± 2.6 pp, medium confidence |
| MMMU-Pro | 88.0% | measured |
| AA-AnalystAgent | 56.3% | measured |
| Harvey LAB | 91.2% | measured |
| Humanity's Last Exam | 61.4% | measured |
| AutomationBench | 69.5% | measured |
| EnterpriseOps-Gym | 54.8% | estimated ± 1.8 pp, low confidence |
| ITBench SRE | 38.2% | measured |