benchgap
Agentic · tools

ExploitGym leaderboard

As of 2026-10-07, the highest measured score on ExploitGym is 42.4% by GPT-6 Astra. 107 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1GPT-6 Astra42.4%measured
2Atria Dawn Preview39.9%estimated ± 6.1 pp, low confidence
3GPT-6.1 Sol35.1%measured
4Claude Opus 534.4%estimated ± 6.1 pp, low confidence
5GPT-5.6 Sol33.7%measured
6GPT-5.5 Pro32.1%estimated ± 6.1 pp, low confidence
7Claude Fable 530.5%estimated ± 10.5 pp, low confidence
8GPT-5.4 Pro29.5%estimated ± 6.1 pp, low confidence
9Qwen3.8 Max Preview26.8%estimated ± 12.2 pp, low confidence
10Claude Mythos 525.3%estimated ± 6.1 pp, low confidence
11Muse Spark 1.324.9%estimated ± 5.0 pp, low confidence
12Claude Fable 5.124.0%estimated ± 5.0 pp, low confidence
13Claude Opus 5.524.0%estimated ± 5.0 pp, low confidence
14GPT-5.6 Terra23.2%measured
15Claude Sonnet 5.523.2%estimated ± 5.0 pp, low confidence
16GPT-6 Sol22.1%measured
17Ornith-1.5-397B20.8%estimated ± 6.1 pp, low confidence
18Muse Spark 1.220.3%estimated ± 12.2 pp, low confidence
19Grok 4.518.2%estimated ± 12.2 pp, low confidence
20Gemini 3.7 Flash18.0%estimated ± 6.9 pp, low confidence
21Ling 3.1 Flash17.8%estimated ± 9.3 pp, low confidence
22Fugu Cyber17.8%estimated ± 9.3 pp, low confidence
23MiMo-V2.6-Pro17.8%measured
24Gemini 3.8 Flash Cyber17.8%estimated ± 9.3 pp, low confidence
25Claude Mythos Preview17.5%measured
26Kimi K317.5%estimated ± 5.0 pp, low confidence
27Gemini 3.5 Flash Cyber17.5%estimated ± 9.3 pp, low confidence
28Gemini 4 Argon17.3%estimated ± 5.0 pp, low confidence
29Gemini 3.8 Flash16.6%estimated ± 5.0 pp, low confidence
30Claude Haiku 5.516.4%estimated ± 5.0 pp, low confidence
31Grok 4.715.8%estimated ± 5.0 pp, low confidence
32Grok 4.615.5%estimated ± 10.5 pp, low confidence
33DeepSeek V4.1 Flash15.3%measured
34Mistral Large 415.0%estimated ± 5.0 pp, low confidence
35GLM-5.315.0%measured
36Claude Sonnet 514.6%estimated ± 6.1 pp, low confidence
37Qwen3.8-27B14.3%estimated ± 5.0 pp, low confidence
38GLM-5.3-Flash14.0%estimated ± 5.0 pp, low confidence
39Step 5 Preview13.9%estimated ± 5.0 pp, low confidence
40Inkling13.7%estimated ± 5.0 pp, low confidence
41Muse Glimmer 30B13.6%estimated ± 5.0 pp, low confidence
42MiniMax M313.6%estimated ± 5.0 pp, low confidence
43Nemotron 3 Ultra13.6%estimated ± 5.0 pp, low confidence
44GLM-5.213.5%estimated ± 10.3 pp, low confidence
45GPT-5.513.4%measured
46Claude Opus 4.813.3%estimated ± 6.1 pp, low confidence
47GPT-5.6 Luna12.4%measured
48GPT-6 Luna11.6%measured
49Claude Opus 4.611.4%estimated ± 6.1 pp, low confidence
50DeepSeek V4 Pro 081310.4%estimated ± 6.1 pp, low confidence
51dots3-note Preview10.1%estimated ± 6.1 pp, low confidence
52Kimi K2.69.8%estimated ± 6.1 pp, low confidence
53Hy4 preview9.7%estimated ± 9.3 pp, low confidence
54Quasar 438B7.3%estimated ± 12.2 pp, low confidence
55GPT-5.46.0%measured
56MiMo-V2.6-Flash6.0%measured
57Qwen3.8-Flash-Next5.2%estimated ± 6.9 pp, low confidence
58Qwen3.8 Max5.2%estimated ± 6.9 pp, low confidence
59Gemini 3.6 Flash4.4%estimated ± 12.2 pp, low confidence
60Claude Opus 4.73.5%estimated ± 6.9 pp, low confidence
61Claude Sonnet 4.62.0%estimated ± 6.9 pp, low confidence
62Claude Opus 4.51.5%estimated ± 9.3 pp, low confidence
63GLM-51.5%estimated ± 9.3 pp, low confidence
64Muse Spark1.5%estimated ± 9.3 pp, low confidence
65Gemini 3.5 Flash1.1%estimated ± 12.2 pp, low confidence
66Muse Spark 1.10.8%measured
67Qwen3.7 Plus0.6%estimated ± 6.9 pp, low confidence
68Agents-A10.0%estimated ± 6.1 pp, low confidence
69Agents-A1-4B0.0%estimated ± 6.1 pp, low confidence
70Celeris-10.0%estimated ± 12.2 pp, low confidence
71Claude Opus 4.7 (Adaptive)0.0%estimated ± 6.1 pp, low confidence
72Command A+0.0%estimated ± 12.2 pp, low confidence
73DeepSeek V30.0%estimated ± 12.2 pp, low confidence
74DeepSeek V4 Flash 07310.0%estimated ± 6.1 pp, low confidence
75Gemini 2.5 Pro0.0%estimated ± 12.2 pp, low confidence
76Gemini 3.1 Pro0.0%estimated ± 12.2 pp, low confidence
77Gemini 3.5 Flash-Lite0.0%estimated ± 12.2 pp, low confidence
78Gemma 3 27B0.0%estimated ± 12.2 pp, low confidence
79Gemma 4 31B0.0%estimated ± 12.2 pp, low confidence
80GLM-4.70.0%estimated ± 6.1 pp, low confidence
81GLM-5.10.0%estimated ± 6.1 pp, low confidence
82GPT-5.20.0%estimated ± 6.1 pp, low confidence
83GPT-5.4 mini0.0%estimated ± 12.2 pp, low confidence
84GPT-5.4 nano0.0%estimated ± 12.2 pp, low confidence
85GPT-OSS 120B0.0%estimated ± 12.2 pp, low confidence
86GPT-OSS 20B0.0%estimated ± 12.2 pp, low confidence
87Granite 4.2 8B0.0%estimated ± 12.2 pp, low confidence
88Grok 4.30.0%estimated ± 12.2 pp, low confidence
89Hy30.0%estimated ± 12.2 pp, low confidence
90Hy3 Preview0.0%estimated ± 12.2 pp, low confidence
91Inkling-Small0.0%estimated ± 6.1 pp, low confidence
92Kimi K2.50.0%estimated ± 6.1 pp, low confidence
93Kimi K2.5 (Reasoning)0.0%estimated ± 6.1 pp, low confidence
94Kimi K2.7 Code0.0%estimated ± 12.2 pp, low confidence
95Ling 3.0 Flash0.0%estimated ± 6.1 pp, low confidence
96Ling 3.0 Flash FP80.0%estimated ± 12.2 pp, low confidence
97Llama 4 Maverick0.0%estimated ± 12.2 pp, low confidence
98Llama 4 Scout0.0%estimated ± 12.2 pp, low confidence
99LongCat-Flash-Lite-Sparse0.0%estimated ± 6.1 pp, low confidence
100MiMo-V2.5-Pro0.0%estimated ± 12.2 pp, low confidence
101MiniMax M2.70.0%estimated ± 12.2 pp, low confidence
102Mistral Large 30.0%estimated ± 12.2 pp, low confidence
103Mistral Medium 3.5 128B0.0%estimated ± 12.2 pp, low confidence
104Mistral Small 40.0%estimated ± 12.2 pp, low confidence
105Mistral Small 4 (Reasoning)0.0%estimated ± 12.2 pp, low confidence
106Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 6.1 pp, low confidence
107Nemotron 3 Nano 30B0.0%estimated ± 12.2 pp, low confidence
108Nemotron 3 Super 100B0.0%estimated ± 12.2 pp, low confidence
109Ornith-1.5-35B-A3B0.0%estimated ± 6.1 pp, low confidence
110Ornith-1.5-9B0.0%estimated ± 6.1 pp, low confidence
111Qwen3.5-122B-A10B0.0%estimated ± 6.1 pp, low confidence
112Qwen3.5-27B0.0%estimated ± 6.1 pp, low confidence
113Qwen3.5-35B-A3B0.0%estimated ± 6.1 pp, low confidence
114Qwen3.5 397B0.0%estimated ± 6.1 pp, low confidence
115Qwen3.6-27B0.0%estimated ± 12.2 pp, low confidence
116Qwen3.6-35B-A3B0.0%estimated ± 12.2 pp, low confidence
117Qwen3.7 Max0.0%estimated ± 12.2 pp, low confidence
118Beam0.0%estimated ± 6.1 pp, low confidence
119Solar Pro 40.0%estimated ± 6.1 pp, low confidence
120Step 3.7 Flash0.0%estimated ± 6.1 pp, low confidence
121Trinity-Large-Preview0.0%estimated ± 12.2 pp, low confidence
122Trinity-Large-Thinking0.0%estimated ± 12.2 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General