Anthropic · model
Claude Fable 5 (with fallback) benchmark scores
As of 2026-10-07, Claude Fable 5 (with fallback) (Anthropic) has measured scores on 7 benchmarks and estimated scores on 4 more.
| Benchmark | Score | Source |
|---|---|---|
| Terminal-Bench Hard | 62.9% | measured |
| GPQA Diamond | 95.1% | estimated ± 1.3 pp, high confidence |
| SciCode | 61.0% | measured |
| CritPt | 30.0% | estimated ± 3.5 pp, high confidence |
| τ³-Bench Banking | 45.3% | estimated ± 4.5 pp, medium confidence |
| AA-AnalystAgent | 48.8% | measured |
| Harvey LAB | 93.6% | measured |
| Humanity's Last Exam | 55.5% | measured |
| AutomationBench | 59.2% | estimated ± 11.2 pp, low confidence |
| EnterpriseOps-Gym | 51.1% | measured |
| τ²-Bench Telecom | 98.5% | measured |