benchgap
Agentic · tools

CyberGym leaderboard

As of 2026-10-07, the highest measured score on CyberGym is 95.1% by MiMo-V2.6-Flash. 151 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.8-Flash-Next100.0%estimated ± 5.4 pp, low confidence
2MiMo-V2.6-Flash95.1%measured
3Claude Opus 5.594.9%estimated ± 2.4 pp, low confidence
4MiMo-V2.6-Pro94.0%measured
5Claude Fable 593.4%estimated ± 5.0 pp, low confidence
6Grok 4.792.1%estimated ± 6.2 pp, low confidence
7Claude Opus 589.1%estimated ± 2.4 pp, low confidence
8Gemini 4 Argon88.8%estimated ± 3.7 pp, high confidence
9Claude Sonnet 5.588.7%estimated ± 2.4 pp, low confidence
10DeepSeek V4.1 Flash88.1%measured
11Muse Spark 1.388.0%estimated ± 3.7 pp, high confidence
12Ling 3.1 Flash87.9%measured
13Fugu Cyber86.9%measured
14Atria Dawn Preview86.5%measured
15Gemini 3.8 Flash Cyber86.2%measured
16Gemini 3.8 Flash85.6%estimated ± 5.0 pp, medium confidence
17Step 5 Preview84.7%measured
18GLM-5.384.5%measured
19GPT-5.6 Sol84.5%measured
20Holo3-35B-A3B84.3%estimated ± 11.3 pp, low confidence
21Claude Mythos 583.8%measured
22Qwen3.8 Max Preview83.6%estimated ± 5.3 pp, medium confidence
23DeepSeek V4 Pro 081383.3%measured
24Gemini 3.5 Flash Cyber83.2%measured
25Claude Mythos Preview83.1%measured
26GPT-5.5 Pro82.8%estimated ± 4.6 pp, high confidence
27GPT-5.4 Pro82.4%estimated ± 4.6 pp, high confidence
28GPT-6.1 Sol82.1%estimated ± 3.7 pp, high confidence
29GPT-5.581.8%measured
30GPT-5.6 Terra81.8%measured
31Qwen3.8-27B81.3%estimated ± 5.3 pp, medium confidence
32GPT-6 Sol80.8%estimated ± 3.7 pp, high confidence
33Claude Haiku 5.580.1%estimated ± 2.4 pp, medium confidence
34Claude Sonnet 580.1%estimated ± 2.4 pp, medium confidence
35Claude Fable 5.180.0%estimated ± 3.7 pp, high confidence
36GPT-6 Astra79.9%estimated ± 2.4 pp, medium confidence
37Kimi K379.7%estimated ± 3.7 pp, high confidence
38Claude Opus 4.879.6%estimated ± 4.6 pp, high confidence
39Gemini 3.7 Flash79.6%estimated ± 3.7 pp, high confidence
40GPT-6 Luna79.5%estimated ± 6.2 pp, medium confidence
41Muse Spark 1.279.3%estimated ± 5.3 pp, medium confidence
42Qwen3.8 Max79.3%estimated ± 2.4 pp, medium confidence
43Apodex 1.179.2%estimated ± 2.4 pp, medium confidence
44Ornith-1.5-397B79.2%estimated ± 2.4 pp, medium confidence
45MiniMax M379.2%estimated ± 4.6 pp, high confidence
46Kimi K2.679.0%estimated ± 4.6 pp, high confidence
47GPT-5.479.0%measured
48GLM-5.3-Flash78.8%estimated ± 2.4 pp, medium confidence
49Holo3-122B-A10B78.8%estimated ± 11.3 pp, low confidence
50Mistral Large 478.8%estimated ± 6.2 pp, medium confidence
51Hy4 preview78.4%measured
52Qwen3.7 Max78.1%estimated ± 2.4 pp, medium confidence
53GPT-5.6 Luna77.9%measured
54Grok 4.577.9%estimated ± 5.3 pp, medium confidence
55dots3-note Preview77.8%estimated ± 2.4 pp, medium confidence
56Grok 4.677.1%estimated ± 5.0 pp, medium confidence
57Agents-A176.8%estimated ± 2.4 pp, medium confidence
58Step 3.7 Flash76.8%estimated ± 2.4 pp, medium confidence
59DeepSeek V4 Flash 073176.7%measured
60Nemotron 3 Ultra76.3%estimated ± 2.4 pp, low confidence
61Ornith-1.5-35B-A3B76.3%estimated ± 2.4 pp, low confidence
62Ornith-1.5-9B76.3%estimated ± 2.4 pp, low confidence
63GLM-5.275.8%estimated ± 5.3 pp, medium confidence
64UI-Mate-27B75.6%estimated ± 11.3 pp, low confidence
65Inkling-Small75.6%estimated ± 4.6 pp, high confidence
66Beam75.6%estimated ± 4.6 pp, high confidence
67Inkling75.4%estimated ± 4.6 pp, high confidence
68Claude Opus 4.7 (Adaptive)73.1%measured
69Claude Opus 4.772.5%estimated ± 5.0 pp, medium confidence
70Ling 3.0 Flash72.3%estimated ± 4.6 pp, high confidence
71Quasar 438B70.7%estimated ± 5.3 pp, medium confidence
72Agents-A1-4B68.8%estimated ± 4.6 pp, medium confidence
73Gemini 3.6 Flash68.7%estimated ± 5.3 pp, medium confidence
74GLM-5.168.7%measured
75GPT-5.268.1%estimated ± 4.6 pp, medium confidence
76Qwen3.5-122B-A10B66.7%estimated ± 4.6 pp, medium confidence
77Claude Opus 4.666.6%measured
78Gemini 3.5 Flash66.5%estimated ± 5.3 pp, medium confidence
79Solar Open 266.1%estimated ± 9.9 pp, medium confidence
80Apodex 1.1 Mini65.8%estimated ± 6.2 pp, medium confidence
81Gemini 3 Pro65.7%estimated ± 10.0 pp, low confidence
82Qwen3.5 397B65.5%estimated ± 4.6 pp, medium confidence
83Hy365.3%estimated ± 5.3 pp, medium confidence
84Hy3 Preview65.3%estimated ± 5.3 pp, medium confidence
85Claude Sonnet 4.665.2%measured
86Qwen3.5-27B64.8%estimated ± 4.6 pp, medium confidence
87Qwen3.5-35B-A3B64.8%estimated ± 4.6 pp, medium confidence
88Kimi K2.564.5%estimated ± 4.6 pp, medium confidence
89Kimi K2.5 (Reasoning)64.5%estimated ± 4.6 pp, medium confidence
90Ling 3.0 Flash VL63.7%estimated ± 6.2 pp, medium confidence
91MiMo-V2.5-Pro63.0%estimated ± 5.3 pp, low confidence
92Kimi K2.7 Code62.8%estimated ± 5.3 pp, low confidence
93Ling 3.0 Flash FP861.7%estimated ± 5.3 pp, low confidence
94Qwen3.6-27B61.0%estimated ± 5.3 pp, low confidence
95Qwen3.7 Plus60.7%estimated ± 5.3 pp, low confidence
96GPT-5.4 mini60.6%estimated ± 5.3 pp, low confidence
97GPT-5.4 nano59.2%estimated ± 5.3 pp, low confidence
98Muse Spark 1.159.0%measured
99Grok 4.358.8%estimated ± 5.3 pp, low confidence
100MiniMax M2.758.5%estimated ± 5.3 pp, low confidence
101LLaDA2.2-flash58.4%estimated ± 9.9 pp, medium confidence
102GLM-4.758.0%estimated ± 4.6 pp, medium confidence
103Gemini 3.5 Flash-Lite57.8%estimated ± 5.3 pp, low confidence
104Qwen3.6-35B-A3B57.1%estimated ± 5.3 pp, low confidence
105GPT-5.3 Codex56.4%estimated ± 10.0 pp, low confidence
106Solar Pro 455.7%estimated ± 4.6 pp, medium confidence
107LongCat-Flash-Lite-Sparse55.2%estimated ± 4.6 pp, medium confidence
108Gemini 3 Flash55.1%estimated ± 10.0 pp, low confidence
109Muse Glimmer 30B53.6%estimated ± 5.3 pp, low confidence
110Gemini 3.1 Pro53.5%estimated ± 5.3 pp, low confidence
111Mistral Medium 3.5 128B52.7%estimated ± 5.3 pp, low confidence
112UI-Mate-9B52.4%estimated ± 11.3 pp, low confidence
113Qwen3.6 Plus51.7%estimated ± 6.2 pp, low confidence
114Gemma 4 31B50.7%estimated ± 5.3 pp, low confidence
115Claude Opus 4.550.6%measured
116GPT-OSS 120B50.3%estimated ± 5.3 pp, low confidence
117A.X K249.3%estimated ± 6.2 pp, low confidence
118Nemotron 3 Super 100B48.7%estimated ± 5.3 pp, low confidence
119Granite 4.2 8B48.4%estimated ± 5.3 pp, low confidence
120Command A+48.3%estimated ± 5.3 pp, low confidence
121Gemini 2.5 Pro48.3%estimated ± 5.3 pp, low confidence
122GPT-5.2-Codex47.9%estimated ± 10.0 pp, low confidence
123Claude 4.1 Opus47.6%estimated ± 11.4 pp, low confidence
124Mistral Large 347.4%estimated ± 5.3 pp, low confidence
125Mistral Small 446.7%estimated ± 5.3 pp, low confidence
126Mistral Small 4 (Reasoning)46.7%estimated ± 5.3 pp, low confidence
127GPT-OSS 20B46.6%estimated ± 5.3 pp, low confidence
128Trinity-Large-Preview46.4%estimated ± 5.3 pp, low confidence
129Trinity-Large-Thinking46.4%estimated ± 5.3 pp, low confidence
130Nemotron 3 Nano 30B46.3%estimated ± 5.3 pp, low confidence
131DeepSeek V346.2%estimated ± 5.3 pp, low confidence
132Celeris-146.1%estimated ± 5.3 pp, low confidence
133Llama 4 Maverick46.0%estimated ± 5.3 pp, low confidence
134Llama 4 Scout46.0%estimated ± 5.3 pp, low confidence
135GPT-5 (high)45.7%estimated ± 6.2 pp, low confidence
136Gemma 3 27B45.7%estimated ± 5.3 pp, low confidence
137GPT-5.1-Codex44.9%estimated ± 10.0 pp, low confidence
138Nemotron 3.5 Lightning 30B A3B NVFP444.8%estimated ± 4.6 pp, medium confidence
139Grok Build 0.144.2%estimated ± 10.0 pp, low confidence
140Muse Spark43.5%measured
141Claude Sonnet 4.543.3%estimated ± 10.0 pp, low confidence
142GLM-543.2%measured
143Qwen3.5 Plus42.2%estimated ± 11.4 pp, low confidence
144Grok 4.1 Fast41.6%estimated ± 10.0 pp, low confidence
145MiMo-V2.541.1%estimated ± 10.0 pp, low confidence
146Claude Haiku 4.537.9%estimated ± 11.4 pp, low confidence
147GPT-5.137.8%estimated ± 6.2 pp, low confidence
148Qwen3 Max36.9%estimated ± 10.0 pp, low confidence
149Grok 435.1%estimated ± 10.0 pp, low confidence
150Claude 4 Sonnet31.7%estimated ± 10.0 pp, low confidence
151Gemini 3.1 Flash-Lite30.2%estimated ± 10.0 pp, low confidence
152Grok 4.2030.1%estimated ± 10.0 pp, low confidence
153MiMo-V2-Pro28.0%estimated ± 10.0 pp, low confidence
154MiniCPM5-2B25.8%estimated ± 6.2 pp, low confidence
155Laguna S 2.122.5%estimated ± 6.0 pp, low confidence
156GLM-5V-Turbo21.1%estimated ± 10.0 pp, low confidence
157DeepSeek V3.219.8%estimated ± 10.0 pp, low confidence
158MiMo-V2-Flash17.1%estimated ± 6.2 pp, low confidence
159GPT-4.115.6%estimated ± 10.0 pp, low confidence
160Granite 4.2 30B10.4%estimated ± 6.2 pp, low confidence
161Gemma 4 26B A4B9.2%estimated ± 6.2 pp, low confidence
162Ling 3.0 Tiny8.7%estimated ± 6.2 pp, low confidence
163DeepSeek V3 03240.0%estimated ± 6.2 pp, low confidence
164Gemma 4 12B0.0%estimated ± 6.2 pp, low confidence
165Gemma 4 E2B0.0%estimated ± 6.2 pp, low confidence
166Gemma 4 E4B0.0%estimated ± 6.2 pp, low confidence
167GPT-4.1 mini0.0%estimated ± 6.2 pp, low confidence
168GPT-4.1 nano0.0%estimated ± 6.2 pp, low confidence
169GPT-4o0.0%estimated ± 6.2 pp, low confidence
170GPT-4o mini0.0%estimated ± 6.2 pp, low confidence
171Granite 4.2 3B0.0%estimated ± 6.2 pp, low confidence
172K-Exaone0.0%estimated ± 6.2 pp, low confidence
173LFM2.5-2.6B0.0%estimated ± 6.2 pp, low confidence
174Ling 2.6 Flash0.0%estimated ± 6.2 pp, low confidence
175Mercury 2.50.0%estimated ± 6.2 pp, low confidence
176Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 6.2 pp, low confidence
177North Mini Code0.0%estimated ± 6.2 pp, low confidence
178Solar Pro 30.0%estimated ± 6.2 pp, low confidence
179Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 6.2 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General