benchgap
General

ExploitBench leaderboard

As of 2026-10-07, the highest measured score on ExploitBench is 100.0% by GPT-6 Astra. 4 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra100.0%measured
2GPT-6.1 Sol99.7%measured
3Claude Mythos 578.0%measured
4GPT-5.6 Sol73.5%measured
5Claude Mythos Preview69.0%measured
6GPT-6 Sol61.1%estimated ± 12.2 pp, low confidence
7GLM-5.354.4%measured
8GPT-5.6 Terra52.9%measured
9DeepSeek V4.1 Flash52.5%estimated ± 12.2 pp, low confidence
10MiMo-V2.6-Pro47.9%measured
11GPT-5.6 Luna33.2%measured
12Kimi K332.0%measured
13MiMo-V2.6-Flash25.3%measured
14GPT-5.520.0%estimated ± 12.2 pp, low confidence
15GPT-6 Luna12.5%estimated ± 12.2 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General