benchgap
Anthropic · model

Claude Fable 5.1 (xhigh with fallback) benchmark scores

As of 2026-10-07, Claude Fable 5.1 (xhigh with fallback) (Anthropic) has measured scores on 6 benchmarks and estimated scores on 6 more.

BenchmarkScoreSource
Terminal-Bench 2.191.0%measured
Terminal-Bench 4.055.1%measured
Terminal-Bench Science 0.149.4%estimated ± 6.5 pp, medium confidence
GPQA Diamond95.6%estimated ± 1.3 pp, high confidence
SciCode60.9%measured
CritPt31.1%measured
τ³-Bench Banking40.5%estimated ± 10.4 pp, low confidence
AA-AnalystAgent47.3%estimated ± 13.9 pp, low confidence
Harvey LAB93.3%measured
Humanity's Last Exam58.7%measured
AutomationBench59.6%estimated ± 13.0 pp, low confidence
EnterpriseOps-Gym46.0%estimated ± 4.4 pp, medium confidence