benchgap
Anthropic · model

Claude Mythos 5 benchmark scores

As of 2026-10-07, Claude Mythos 5 (Anthropic) has measured scores on 13 benchmarks and estimated scores on 82 more.

BenchmarkScoreSource
AA-GPQA Diamond93.6%estimated ± 2.7 pp, high confidence
AA-HLE55.2%estimated ± 5.0 pp, high confidence
AA-SciCode61.0%estimated ± 3.8 pp, high confidence
Artificial Analysis Intelligence Index51.5%estimated ± 5.1 pp, medium confidence
CritPt24.6%estimated ± 4.1 pp, high confidence
GDPval-AA51.1%estimated ± 5.1 pp, medium confidence
GPQA Diamond94.1%estimated ± 0.4 pp, high confidence
HLE64.5%measured
LiveCodeBench v696.2%estimated ± 2.1 pp, high confidence
AA-Omniscience Accuracy51.2%estimated ± 8.1 pp, medium confidence
SciCode55.3%estimated ± 3.8 pp, medium confidence
Terminal-Bench 2.188.0%measured
BrowseComp88.0%measured
HLE w/ tools58.8%estimated ± 4.6 pp, high confidence
VITA-Bench46.1%estimated ± 9.9 pp, low confidence
AA Coding Index78.9%estimated ± 5.1 pp, medium confidence
AA-MMMU-Pro82.9%estimated ± 3.2 pp, medium confidence
APEX-Agents-AA37.8%estimated ± 4.6 pp, high confidence
DeepSearchQA90.6%estimated ± 8.1 pp, low confidence
SWE-bench Verified95.5%measured
AutomationBench36.9%estimated ± 5.8 pp, medium confidence
CyberGym83.8%measured
JobBench52.4%estimated ± 8.6 pp, low confidence
SWE-bench Pro80.3%measured
WideResearch81.2%estimated ± 4.6 pp, high confidence
AA Agentic Index48.5%estimated ± 5.8 pp, medium confidence
MMLU-Pro86.0%estimated ± 2.5 pp, low confidence
GPQA94.1%measured
Gert Labs69.3%estimated ± 6.0 pp, low confidence
AA-AnalystAgent39.6%estimated ± 12.5 pp, low confidence
AA EnterpriseOps-Gym49.9%estimated ± 5.9 pp, low confidence
AA Harvey LAB92.4%estimated ± 9.8 pp, low confidence
ApprenticeBench24.2%estimated ± 4.0 pp, medium confidence
ARC-AGI-197.4%estimated ± 3.7 pp, high confidence
ARC-AGI-279.5%estimated ± 11.6 pp, low confidence
CursorBench 3.166.9%estimated ± 5.6 pp, low confidence
CursorBench 3.271.4%estimated ± 3.6 pp, high confidence
Terminal-Bench 3.023.0%estimated ± 9.1 pp, low confidence
FrontierCode 1.1 Main53.2%estimated ± 2.7 pp, medium confidence
FrontierSWE v251.3%estimated ± 9.0 pp, medium confidence
OSWorld-Verified85.0%measured
Vals GPQA Diamond94.4%estimated ± 3.0 pp, high confidence
Vals LiveCodeBench88.5%estimated ± 4.7 pp, medium confidence
Vals MMLU-Pro88.5%estimated ± 1.2 pp, high confidence
Vals SWE-bench97.7%estimated ± 6.8 pp, medium confidence
Terminal-Bench 2.1 (Vals AI)73.9%estimated ± 7.4 pp, medium confidence
AA AutomationBench46.4%estimated ± 5.7 pp, low confidence
GDP.pdf12.2%estimated ± 6.6 pp, low confidence
AA Tau3 Banking35.9%estimated ± 10.2 pp, low confidence
AA Terminal-Bench 2.184.9%estimated ± 3.5 pp, high confidence
AA Terminal-Bench 4.028.6%estimated ± 8.9 pp, medium confidence
CursorBench 4.051.0%estimated ± 3.6 pp, high confidence
HLE w/o tools59.0%measured
OSWorld 2.020.4%estimated ± 3.6 pp, high confidence
PostTrainBench v1.140.1%estimated ± 4.7 pp, high confidence
SWE Multilingual92.2%measured
Toolathlon-Verified73.5%estimated ± 2.2 pp, low confidence
Vibe Code Bench62.5%estimated ± 8.1 pp, low confidence
CharXiv93.5%measured
CharXiv w/o tools88.9%measured
ExploitBench78.0%measured
ExploitGym25.3%estimated ± 6.1 pp, low confidence
C-Eval93.3%estimated ± 0.8 pp, low confidence
Claw-Eval81.7%estimated ± 7.9 pp, low confidence
MathVision96.3%estimated ± 2.1 pp, low confidence
MCP Atlas77.1%estimated ± 10.0 pp, low confidence
MMLU-ProX91.1%estimated ± 1.7 pp, low confidence
MMLU-Redux97.3%estimated ± 2.1 pp, low confidence
MMMU-Pro82.3%estimated ± 1.4 pp, low confidence
NL2Repo76.7%estimated ± 6.1 pp, low confidence
NOVA-6363.9%estimated ± 1.6 pp, low confidence
ScreenSpot Pro89.8%estimated ± 7.1 pp, medium confidence
SuperGPQA87.6%estimated ± 9.1 pp, low confidence
Toolathlon55.9%estimated ± 3.5 pp, high confidence
VideoMMMU88.2%estimated ± 1.5 pp, high confidence
ERQA73.0%estimated ± 3.7 pp, high confidence
HealthBench Hard36.4%estimated ± 5.1 pp, low confidence
LiveCodeBench Pro97.4%estimated ± 6.0 pp, low confidence
MedXpertQA (MM)80.4%estimated ± 6.4 pp, low confidence
MedXpertQA (Text)74.6%estimated ± 8.9 pp, low confidence
ResearchClawBench19.3%estimated ± 2.2 pp, low confidence
SWE-Rebench61.8%estimated ± 3.6 pp, low confidence
ARC-AGI-32.8%estimated ± 2.1 pp, high confidence
FrontierCode 1.1 Extended61.5%estimated ± 3.0 pp, medium confidence
HealthBench Professional58.1%estimated ± 4.8 pp, medium confidence
LiveCodeBench95.3%estimated ± 10.2 pp, low confidence
Agents' Last Exam27.0%estimated ± 2.1 pp, high confidence
OpenHarmony Bench60.9%estimated ± 4.9 pp, low confidence
SEC-Bench Pro76.1%estimated ± 4.3 pp, medium confidence
SimpleVQA73.6%estimated ± 7.8 pp, low confidence
MMLU90.5%estimated ± 3.6 pp, low confidence
MathVision w/ Python97.7%estimated ± 0.9 pp, medium confidence
RealWorldQA88.0%estimated ± 1.1 pp, medium confidence
Video-MME (with subtitle)90.1%estimated ± 0.9 pp, medium confidence
OmniDocBench 1.592.7%estimated ± 7.4 pp, low confidence