benchgap
Anthropic · model

Claude Mythos Preview benchmark scores

As of 2026-10-07, Claude Mythos Preview (Anthropic) has measured scores on 3 benchmarks and estimated scores on 20 more.

BenchmarkScoreSource
GDPval-AA50.5%estimated ± 5.1 pp, medium confidence
BrowseComp85.7%estimated ± 1.8 pp, medium confidence
HLE w/ tools59.1%estimated ± 5.0 pp, low confidence
APEX-Agents-AA35.8%estimated ± 6.0 pp, low confidence
DeepSearchQA76.8%estimated ± 11.5 pp, low confidence
AutomationBench31.3%estimated ± 5.8 pp, medium confidence
CyberGym83.1%measured
JobBench52.1%estimated ± 8.6 pp, low confidence
AA Agentic Index47.5%estimated ± 5.8 pp, medium confidence
Gert Labs69.1%estimated ± 6.0 pp, low confidence
ApprenticeBench22.2%estimated ± 4.0 pp, medium confidence
OSWorld-Verified79.7%estimated ± 4.7 pp, high confidence
AA AutomationBench61.1%estimated ± 5.5 pp, low confidence
GDP.pdf19.3%estimated ± 5.9 pp, low confidence
AA ITBench46.3%estimated ± 3.9 pp, medium confidence
OSWorld 2.040.2%estimated ± 13.6 pp, low confidence
Toolathlon-Verified73.4%estimated ± 2.2 pp, low confidence
ExploitBench69.0%measured
ExploitGym17.5%measured
MCP Atlas78.5%estimated ± 11.4 pp, low confidence
Toolathlon54.9%estimated ± 9.0 pp, low confidence
Agents' Last Exam25.2%estimated ± 2.1 pp, high confidence
SEC-Bench Pro73.3%estimated ± 4.3 pp, medium confidence