benchgap
Anthropic · model

Claude Fable 5.1 benchmark scores

As of 2026-10-07, Claude Fable 5.1 (Anthropic) has measured scores on 41 benchmarks and estimated scores on 61 more.

BenchmarkScoreSource
AA-GPQA Diamond93.7%measured
AA-HLE59.1%measured
AA-SciCode63.1%measured
Artificial Analysis Intelligence Index53.3%measured
CritPt29.7%measured
GDPval-AA62.9%measured
GPQA Diamond94.7%estimated ± 1.4 pp, high confidence
HLE65.0%measured
AA-LCR85.3%measured
LiveCodeBench v698.9%estimated ± 2.1 pp, medium confidence
AA-Omniscience Accuracy67.2%measured
SciCode57.8%estimated ± 2.9 pp, low confidence
Terminal-Bench 2.189.5%estimated ± 3.9 pp, medium confidence
BrowseComp84.5%estimated ± 2.2 pp, high confidence
HLE w/ tools58.2%estimated ± 4.6 pp, medium confidence
LongBench v261.5%estimated ± 5.9 pp, low confidence
τ²-bench results85.4%estimated ± 4.8 pp, medium confidence
AA Coding Index81.6%measured
APEX-Agents49.1%estimated ± 10.0 pp, low confidence
APEX-Agents-AA38.2%estimated ± 2.9 pp, medium confidence
BioMysteryBench (human-difficult)51.6%estimated ± 5.9 pp, low confidence
DeepSearchQA94.3%estimated ± 4.5 pp, medium confidence
SWE-bench Verified94.6%estimated ± 2.5 pp, medium confidence
AutomationBench31.4%measured
BFCL v478.3%estimated ± 2.6 pp, medium confidence
CyberGym80.0%estimated ± 3.7 pp, high confidence
JobBench65.2%estimated ± 3.4 pp, low confidence
SWE-bench Pro81.2%measured
WideResearch85.6%estimated ± 3.9 pp, low confidence
AA Agentic Index58.0%measured
MMLU-Pro86.0%estimated ± 2.5 pp, low confidence
GPQA95.9%estimated ± 1.6 pp, medium confidence
Gert Labs100.0%estimated ± 3.0 pp, low confidence
AA-AnalystAgent57.5%measured
AA EnterpriseOps-Gym56.3%estimated ± 1.8 pp, low confidence
AA Harvey LAB93.0%measured
ApprenticeBench72.0%measured
ARC-AGI-197.5%measured
ARC-AGI-290.0%measured
CursorBench 3.176.3%estimated ± 3.4 pp, low confidence
CursorBench 3.273.4%measured
Terminal-Bench 3.034.5%estimated ± 6.9 pp, medium confidence
FrontierCode 1.1 Main57.9%estimated ± 1.6 pp, low confidence
FrontierSWE v256.3%measured
OSWorld-Verified100.0%estimated ± 1.9 pp, low confidence
Vals GPQA Diamond93.4%measured
Vals LiveCodeBench90.5%measured
Vals MMLU-Pro92.4%measured
Vals SWE-bench100.0%estimated ± 4.5 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)85.0%measured
VulcanBench v387.0%estimated ± 0.7 pp, low confidence
AA AutomationBench59.4%measured
GDP.pdf26.2%measured
AA ITBench49.5%measured
AA Tau3 Banking47.2%measured
AA Terminal-Bench 2.191.4%measured
AA Terminal-Bench 4.052.0%measured
CursorBench 4.051.8%measured
CWE-bench v158.0%measured
DeepSWE67.4%measured
HLE w/o tools60.9%measured
OSWorld 2.041.7%measured
PostTrainBench v1.140.2%measured
ProgramBench87.6%measured
SWE Multilingual89.1%measured
Terminal-Bench 4.055.8%measured
Terminal-Bench-Science 0.152.6%measured
Toolathlon-Verified77.8%measured
Vibe Code Bench100.0%estimated ± 10.3 pp, low confidence
ExploitGym24.0%estimated ± 5.0 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval86.8%estimated ± 13.6 pp, low confidence
MCP Atlas85.3%estimated ± 2.7 pp, low confidence
MMLU-ProX90.0%estimated ± 1.7 pp, low confidence
MMLU-Redux98.1%estimated ± 1.1 pp, low confidence
NL2Repo60.7%estimated ± 5.1 pp, low confidence
NOVA-6362.8%estimated ± 1.6 pp, low confidence
QwenClawBench100.0%estimated ± 6.0 pp, low confidence
SuperGPQA90.9%estimated ± 7.5 pp, low confidence
Terminal-Bench 2.082.7%estimated ± 5.9 pp, low confidence
Toolathlon67.2%estimated ± 1.7 pp, low confidence
HealthBench Hard35.0%estimated ± 7.7 pp, low confidence
LiveCodeBench Pro97.9%estimated ± 6.0 pp, low confidence
MedXpertQA (Text)76.1%estimated ± 8.9 pp, low confidence
MMLU-Pro (Arcee)100.0%estimated ± 4.0 pp, low confidence
React Native Evals84.6%estimated ± 3.5 pp, medium confidence
ResearchClawBench21.5%estimated ± 2.0 pp, medium confidence
SWE-Rebench58.1%estimated ± 8.9 pp, low confidence
ARC-AGI-373.5%estimated ± 3.4 pp, medium confidence
BioMysteryBench (human-solvable)89.5%estimated ± 0.6 pp, medium confidence
DRACO87.3%estimated ± 3.0 pp, medium confidence
FrontierCode 1.1 Extended62.5%estimated ± 1.6 pp, low confidence
HealthBench (raw)64.9%estimated ± 6.4 pp, low confidence
HealthBench (length-adjusted)58.0%estimated ± 2.5 pp, low confidence
HealthBench Professional60.0%estimated ± 1.3 pp, low confidence
HealthBench Professional (raw)74.3%estimated ± 5.0 pp, medium confidence
HLE-Verified54.6%estimated ± 0.7 pp, medium confidence
LABBench284.1%estimated ± 1.3 pp, medium confidence
LiveCodeBench100.0%estimated ± 8.9 pp, low confidence
Agents' Last Exam50.8%estimated ± 6.3 pp, medium confidence
OpenHarmony Bench58.8%estimated ± 2.4 pp, medium confidence
MMLU96.6%estimated ± 1.1 pp, low confidence