benchgap
Anthropic · model

Claude Mythos 5.1 benchmark scores

As of 2026-10-07, Claude Mythos 5.1 (Anthropic) has measured scores on 1 benchmark and estimated scores on 2 more.

BenchmarkScoreSource
Terminal-Bench 2.1 (Vals AI)89.3%estimated ± 4.6 pp, low confidence
AA Terminal-Bench 4.056.5%estimated ± 4.7 pp, high confidence
Terminal-Bench 4.060.9%measured