Anthropic · model
Claude Mythos Preview benchmark scores
As of 2026-10-07, Claude Mythos Preview (Anthropic) has measured scores on 3 benchmarks and estimated scores on 20 more.
| Benchmark | Score | Source |
|---|---|---|
| GDPval-AA | 50.5% | estimated ± 5.1 pp, medium confidence |
| BrowseComp | 85.7% | estimated ± 1.8 pp, medium confidence |
| HLE w/ tools | 59.1% | estimated ± 5.0 pp, low confidence |
| APEX-Agents-AA | 35.8% | estimated ± 6.0 pp, low confidence |
| DeepSearchQA | 76.8% | estimated ± 11.5 pp, low confidence |
| AutomationBench | 31.3% | estimated ± 5.8 pp, medium confidence |
| CyberGym | 83.1% | measured |
| JobBench | 52.1% | estimated ± 8.6 pp, low confidence |
| AA Agentic Index | 47.5% | estimated ± 5.8 pp, medium confidence |
| Gert Labs | 69.1% | estimated ± 6.0 pp, low confidence |
| ApprenticeBench | 22.2% | estimated ± 4.0 pp, medium confidence |
| OSWorld-Verified | 79.7% | estimated ± 4.7 pp, high confidence |
| AA AutomationBench | 61.1% | estimated ± 5.5 pp, low confidence |
| GDP.pdf | 19.3% | estimated ± 5.9 pp, low confidence |
| AA ITBench | 46.3% | estimated ± 3.9 pp, medium confidence |
| OSWorld 2.0 | 40.2% | estimated ± 13.6 pp, low confidence |
| Toolathlon-Verified | 73.4% | estimated ± 2.2 pp, low confidence |
| ExploitBench | 69.0% | measured |
| ExploitGym | 17.5% | measured |
| MCP Atlas | 78.5% | estimated ± 11.4 pp, low confidence |
| Toolathlon | 54.9% | estimated ± 9.0 pp, low confidence |
| Agents' Last Exam | 25.2% | estimated ± 2.1 pp, high confidence |
| SEC-Bench Pro | 73.3% | estimated ± 4.3 pp, medium confidence |