benchgap
Anthropic · model

Claude Fable 5 (with fallback) benchmark scores

As of 2026-10-07, Claude Fable 5 (with fallback) (Anthropic) has measured scores on 7 benchmarks and estimated scores on 4 more.

BenchmarkScoreSource
Terminal-Bench Hard62.9%measured
GPQA Diamond95.1%estimated ± 1.3 pp, high confidence
SciCode61.0%measured
CritPt30.0%estimated ± 3.5 pp, high confidence
τ³-Bench Banking45.3%estimated ± 4.5 pp, medium confidence
AA-AnalystAgent48.8%measured
Harvey LAB93.6%measured
Humanity's Last Exam55.5%measured
AutomationBench59.2%estimated ± 11.2 pp, low confidence
EnterpriseOps-Gym51.1%measured
τ²-Bench Telecom98.5%measured