benchgap
Anthropic · model

Claude Fable 5.1 (max with fallback) benchmark scores

As of 2026-10-07, Claude Fable 5.1 (max with fallback) (Anthropic) has measured scores on 13 benchmarks and estimated scores on 2 more.

BenchmarkScoreSource
Terminal-Bench 2.191.4%measured
Terminal-Bench 4.052.0%measured
Terminal-Bench Science 0.143.3%measured
GPQA Diamond93.7%measured
SciCode63.1%measured
CritPt29.7%measured
GDP.pdf26.2%measured
τ³-Bench Banking47.2%measured
MMMU-Pro84.8%estimated ± 2.6 pp, high confidence
AA-AnalystAgent57.5%measured
Harvey LAB93.0%measured
Humanity's Last Exam59.1%measured
AutomationBench59.4%measured
EnterpriseOps-Gym55.4%estimated ± 1.8 pp, low confidence
ITBench SRE49.5%measured