benchgap
Anthropic · model

Claude Fable 5.1 (high with fallback) benchmark scores

As of 2026-10-07, Claude Fable 5.1 (high with fallback) (Anthropic) has measured scores on 3 benchmarks and estimated scores on 7 more.

BenchmarkScoreSource
Terminal-Bench 2.189.9%measured
Terminal-Bench 4.047.8%estimated ± 11.6 pp, low confidence
GPQA Diamond95.2%estimated ± 1.3 pp, high confidence
CritPt30.2%estimated ± 3.5 pp, high confidence
τ³-Bench Banking39.9%estimated ± 10.4 pp, low confidence
AA-AnalystAgent46.1%estimated ± 13.9 pp, low confidence
Harvey LAB93.0%measured
Humanity's Last Exam55.9%measured
AutomationBench58.1%estimated ± 13.0 pp, low confidence
EnterpriseOps-Gym45.5%estimated ± 4.4 pp, medium confidence