Anthropic · model
Claude Fable 5.1 (xhigh with fallback) benchmark scores
As of 2026-10-07, Claude Fable 5.1 (xhigh with fallback) (Anthropic) has measured scores on 6 benchmarks and estimated scores on 6 more.
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench 2.1 | 91.0% | measured |
| Terminal-Bench 4.0 | 55.1% | measured |
| Terminal-Bench Science 0.1 | 49.4% | estimated ± 6.5 pp, medium confidence |
| GPQA Diamond | 95.6% | estimated ± 1.3 pp, high confidence |
| SciCode | 60.9% | measured |
| CritPt | 31.1% | measured |
| τ³-Bench Banking | 40.5% | estimated ± 10.4 pp, low confidence |
| AA-AnalystAgent | 47.3% | estimated ± 13.9 pp, low confidence |
| Harvey LAB | 93.3% | measured |
| Humanity's Last Exam | 58.7% | measured |
| AutomationBench | 59.6% | estimated ± 13.0 pp, low confidence |
| EnterpriseOps-Gym | 46.0% | estimated ± 4.4 pp, medium confidence |