benchgap
Agentic · tools

AA EnterpriseOps-Gym leaderboard

As of 2026-10-07, the highest measured score on AA EnterpriseOps-Gym is 51.1% by Claude Fable 5. 196 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 3.7 Flash57.6%estimated ± 1.8 pp, low confidence
2Claude Fable 5.156.3%estimated ± 1.8 pp, low confidence
3Claude Sonnet 5.556.3%estimated ± 1.8 pp, low confidence
4Claude Opus 5.555.6%estimated ± 1.8 pp, low confidence
5Claude Opus 554.3%estimated ± 1.8 pp, low confidence
6GLM-5.253.8%estimated ± 3.6 pp, low confidence
7GPT-5.453.5%estimated ± 3.6 pp, low confidence
8GPT-6 Astra53.0%estimated ± 1.8 pp, low confidence
9GLM-5-Turbo52.8%estimated ± 3.6 pp, medium confidence
10GLM-5V-Turbo52.8%estimated ± 3.6 pp, medium confidence
11Step 3.7 Flash52.8%estimated ± 3.6 pp, medium confidence
12GPT-5.552.4%estimated ± 1.8 pp, low confidence
13GPT-6.1 Sol52.4%estimated ± 1.8 pp, low confidence
14GLM-552.2%estimated ± 3.6 pp, medium confidence
15Muse Spark 1.351.9%estimated ± 4.4 pp, medium confidence
16GLM-5.151.4%estimated ± 3.6 pp, medium confidence
17Grok 4.351.4%estimated ± 3.6 pp, medium confidence
18Qwen3.6 Plus51.4%estimated ± 3.6 pp, medium confidence
19Claude Fable 551.1%measured
20GPT-5.6 Sol51.0%estimated ± 1.8 pp, medium confidence
21GPT-6 Sol50.4%estimated ± 4.4 pp, medium confidence
22Gemini 3.5 Flash50.1%measured
23Claude Mythos 549.9%estimated ± 5.9 pp, low confidence
24DeepSeek V4 Pro 081349.6%measured
25DeepSeek V4 Flash 073149.4%estimated ± 5.1 pp, low confidence
26Qwen3.8 Max49.1%estimated ± 5.1 pp, low confidence
27Grok 4.648.3%measured
28Hy4 preview48.3%estimated ± 5.1 pp, low confidence
29GLM-4.748.1%estimated ± 3.6 pp, medium confidence
30Kimi K2.648.1%estimated ± 3.6 pp, medium confidence
31Kimi K2.548.1%estimated ± 3.6 pp, medium confidence
32Kimi K2.5 (Reasoning)48.1%estimated ± 3.6 pp, medium confidence
33Qwen 3.6 Max (preview)48.1%estimated ± 3.6 pp, medium confidence
34GPT-6 Luna47.9%estimated ± 4.4 pp, medium confidence
35Gemini 3.6 Flash47.9%estimated ± 5.9 pp, low confidence
36Gemini 3.1 Pro47.6%estimated ± 3.6 pp, medium confidence
37Holo3-35B-A3B47.4%estimated ± 5.9 pp, low confidence
38Qwen3.8 Max Preview47.3%estimated ± 6.7 pp, low confidence
39Step 5 Preview47.2%measured
40Qwen3.6-35B-A3B47.1%estimated ± 3.6 pp, medium confidence
41Qwen3.8-Flash-Next47.0%estimated ± 6.7 pp, low confidence
42Gemini 4 Argon46.9%estimated ± 4.4 pp, high confidence
43MiMo-V2-Pro46.6%estimated ± 3.6 pp, medium confidence
44Claude Sonnet 546.1%estimated ± 5.9 pp, low confidence
45Gemini 3.8 Flash46.1%estimated ± 4.4 pp, high confidence
46Qwen3.7 Max46.0%estimated ± 3.6 pp, medium confidence
47Claude Haiku 5.545.8%estimated ± 4.4 pp, high confidence
48Muse Spark 1.145.7%estimated ± 5.9 pp, low confidence
49Claude Opus 4.845.5%estimated ± 3.6 pp, medium confidence
50Kimi K345.3%measured
51MiMo-V2.5-Pro45.1%estimated ± 3.6 pp, medium confidence
52Mistral Medium 3.5 128B45.1%estimated ± 3.6 pp, medium confidence
53Qwen3.6-27B45.1%estimated ± 3.6 pp, medium confidence
54Grok 4.745.0%estimated ± 4.4 pp, high confidence
55Muse Spark 1.244.8%estimated ± 6.7 pp, low confidence
56Qwen3.5-27B44.6%estimated ± 3.6 pp, medium confidence
57GPT-5.6 Luna44.5%estimated ± 6.7 pp, low confidence
58GPT-5.5 Pro44.3%estimated ± 6.9 pp, low confidence
59Qwen3.8-27B44.2%measured
60MiMo-V2.6-Pro44.2%estimated ± 4.4 pp, high confidence
61Qwen3.5-122B-A10B44.1%estimated ± 3.6 pp, medium confidence
62GPT-5.4 Pro44.1%estimated ± 6.9 pp, low confidence
63Holo3-122B-A10B43.9%estimated ± 5.9 pp, low confidence
64GPT-5.4 mini43.7%estimated ± 3.6 pp, medium confidence
65Grok 4.1 Fast (Reasoning)43.6%estimated ± 3.6 pp, medium confidence
66Mistral Large 443.6%estimated ± 4.4 pp, high confidence
67Ornith-1.5-397B43.2%estimated ± 6.9 pp, low confidence
68Qwen3.7 Plus43.0%estimated ± 3.6 pp, medium confidence
69Grok 4.542.7%estimated ± 6.7 pp, low confidence
70UI-Mate-27B42.2%estimated ± 5.9 pp, low confidence
71GPT-5.4 nano42.2%estimated ± 3.6 pp, medium confidence
72dots3-note Preview42.1%estimated ± 6.9 pp, low confidence
73Claude Opus 4.6 (Adaptive)41.5%estimated ± 3.6 pp, medium confidence
74GPT-5.2-Codex41.5%estimated ± 3.6 pp, medium confidence
75Muse Spark40.4%estimated ± 3.6 pp, medium confidence
76Beam40.1%estimated ± 6.9 pp, low confidence
77MiMo-V2-Omni39.9%estimated ± 3.6 pp, medium confidence
78Gemini 3.5 Flash-Lite39.6%estimated ± 5.9 pp, low confidence
79Agents-A139.5%estimated ± 6.9 pp, low confidence
80Kimi K2.7 Code38.0%estimated ± 3.6 pp, medium confidence
81Trinity-Large-Preview38.0%estimated ± 3.6 pp, medium confidence
82Trinity-Large-Thinking38.0%estimated ± 3.6 pp, medium confidence
83Inkling38.0%measured
84Hy3 Preview37.5%estimated ± 6.7 pp, low confidence
85DeepSeek V4.1 Flash37.5%estimated ± 4.4 pp, high confidence
86Ling 3.0 Flash37.1%estimated ± 5.3 pp, medium confidence
87Claude Opus 4.5 Thinking37.0%estimated ± 3.6 pp, medium confidence
88Apodex 1.137.0%estimated ± 6.7 pp, low confidence
89Apodex 1.1 Mini37.0%estimated ± 6.7 pp, low confidence
90Ornith-1.5-35B-A3B36.8%estimated ± 6.9 pp, low confidence
91Quasar 438B36.6%estimated ± 6.7 pp, low confidence
92Qwen3.5-35B-A3B36.5%estimated ± 3.6 pp, medium confidence
93GLM-5.336.4%measured
94Ling 3.0 Flash VL36.1%estimated ± 6.7 pp, low confidence
95Claude Opus 4.7 (Adaptive)35.5%estimated ± 3.6 pp, medium confidence
96Inkling-Small35.2%estimated ± 6.7 pp, low confidence
97Solar Pro 434.9%estimated ± 6.7 pp, low confidence
98Muse Glimmer 30B34.7%measured
99LFM2.5-8B-A1B34.6%estimated ± 3.6 pp, medium confidence
100Hy334.1%estimated ± 6.7 pp, low confidence
101UI-Mate-9B33.4%estimated ± 5.9 pp, low confidence
102GLM-5.3-Flash33.2%measured
103Gemini 3 Pro33.1%estimated ± 3.6 pp, medium confidence
104A.X K233.0%estimated ± 6.7 pp, low confidence
105Ling 3.0 Flash FP832.9%estimated ± 6.7 pp, low confidence
106Ornith-1.5-9B32.7%estimated ± 6.9 pp, low confidence
107MiniCPM5-2B32.5%estimated ± 6.7 pp, low confidence
108Nemotron 3.5 Lightning 30B A3B NVFP432.5%estimated ± 6.7 pp, low confidence
109Celeris-132.5%estimated ± 6.7 pp, low confidence
110GPT-4o mini32.5%estimated ± 6.7 pp, low confidence
111Granite 4.2 30B32.5%estimated ± 6.7 pp, low confidence
112Granite 4.2 3B32.5%estimated ± 6.7 pp, low confidence
113Granite 4.2 8B32.5%estimated ± 6.7 pp, low confidence
114LFM2.5-2.6B32.5%estimated ± 6.7 pp, low confidence
115Ling 3.0 Tiny32.5%estimated ± 6.7 pp, low confidence
116Mercury 2.532.5%estimated ± 6.7 pp, low confidence
117GPT-5 (medium)32.1%estimated ± 3.6 pp, medium confidence
118MiniMax M332.1%measured
119MiMo-V2.6-Flash31.9%estimated ± 5.1 pp, low confidence
120Claude Opus 4.531.8%estimated ± 3.6 pp, medium confidence
121GPT-5.6 Terra31.8%estimated ± 3.6 pp, medium confidence
122Solar Pro 331.8%estimated ± 3.6 pp, medium confidence
123Ling 3.1 Flash31.7%estimated ± 5.1 pp, low confidence
124GPT-5.3 Codex31.3%estimated ± 3.6 pp, medium confidence
125Ling 2.6 Flash31.3%estimated ± 3.6 pp, medium confidence
126Atria Dawn Preview30.0%estimated ± 5.1 pp, low confidence
127Claude Sonnet 4.529.9%estimated ± 5.9 pp, low confidence
128LongCat-Flash-Lite-Sparse29.8%estimated ± 6.9 pp, low confidence
129Command A+29.8%estimated ± 3.6 pp, medium confidence
130Claude Opus 4.629.5%estimated ± 3.6 pp, medium confidence
131GPT-5.229.5%estimated ± 3.6 pp, medium confidence
132GPT-5 (high)29.5%estimated ± 3.6 pp, medium confidence
133MiniMax M2.729.5%estimated ± 3.6 pp, medium confidence
134Nemotron 3 Ultra28.9%measured
135MiMo-V2-Flash28.1%estimated ± 3.6 pp, medium confidence
136Qwen3.5 397B28.1%estimated ± 3.6 pp, medium confidence
137Qwen3.5 397B (Reasoning)28.1%estimated ± 3.6 pp, medium confidence
138GPT-5.1-Codex26.8%estimated ± 3.6 pp, low confidence
139GPT-5.1-Codex-Max26.8%estimated ± 3.6 pp, low confidence
140GPT-5.125.2%estimated ± 3.6 pp, low confidence
141o323.6%estimated ± 3.6 pp, low confidence
142LLaDA2.2-flash23.1%estimated ± 3.6 pp, low confidence
143Ternary Bonsai 2 27B22.9%estimated ± 3.6 pp, low confidence
144Claude Sonnet 4.622.0%estimated ± 3.6 pp, low confidence
145DeepSeek V3.221.2%estimated ± 3.6 pp, low confidence
146Agents-A1-4B20.3%estimated ± 3.6 pp, low confidence
147GLM-4.618.8%estimated ± 3.6 pp, low confidence
148Grok Code Fast 117.4%estimated ± 3.6 pp, low confidence
149Grok 416.5%estimated ± 3.6 pp, low confidence
150K-Exaone15.9%estimated ± 3.6 pp, low confidence
151Qwen3 Max15.9%estimated ± 3.6 pp, low confidence
152Claude Opus 4.715.6%estimated ± 3.6 pp, low confidence
153Claude 4.1 Opus Thinking13.1%estimated ± 3.6 pp, low confidence
154Nemotron 3 Super 100B10.1%estimated ± 3.6 pp, low confidence
155GPT-OSS 120B8.7%estimated ± 3.6 pp, low confidence
156Grok 4 Fast (Reasoning)8.7%estimated ± 3.6 pp, low confidence
157Grok 4.1 Fast7.5%estimated ± 3.6 pp, low confidence
158o16.9%estimated ± 3.6 pp, low confidence
159Kimi K26.1%estimated ± 3.6 pp, low confidence
160GPT-OSS 20B5.7%estimated ± 3.6 pp, low confidence
161Gemma 4 31B5.6%estimated ± 3.6 pp, low confidence
162LLaDA2.2-mini4.7%estimated ± 3.6 pp, low confidence
163Gemini 2.5 Pro3.6%estimated ± 3.6 pp, low confidence
164GPT-4.1 mini3.3%estimated ± 3.6 pp, low confidence
165Claude 4 Sonnet3.2%estimated ± 3.6 pp, low confidence
166DeepSeek V3 03242.2%estimated ± 3.6 pp, low confidence
167GPT-4.12.2%estimated ± 3.6 pp, low confidence
168Sarvam 105B2.2%estimated ± 3.6 pp, low confidence
169GLM-4.5-Air2.2%estimated ± 3.6 pp, low confidence
170Nemotron 3 Nano Omni 30B A3B2.0%estimated ± 3.6 pp, low confidence
171Gemma 4 26B A4B1.8%estimated ± 3.6 pp, low confidence
172Gemini 3 Flash1.8%estimated ± 3.6 pp, low confidence
173Mistral Small 41.6%estimated ± 3.6 pp, low confidence
174Mistral Small 4 (Reasoning)1.6%estimated ± 3.6 pp, low confidence
175Nemotron 3 Nano 30B1.6%estimated ± 3.6 pp, low confidence
176DeepSeek V3.1 (Reasoning)1.4%estimated ± 3.6 pp, low confidence
177North Mini Code1.4%estimated ± 3.6 pp, low confidence
178DeepSeek-R11.4%estimated ± 3.6 pp, low confidence
179Gemma 4 12B1.4%estimated ± 3.6 pp, low confidence
180DeepSeek V3.11.3%estimated ± 3.6 pp, low confidence
181Sarvam 30B1.3%estimated ± 3.6 pp, low confidence
182Solar Pro 21.3%estimated ± 3.6 pp, low confidence
183Mistral Large 21.2%estimated ± 3.6 pp, low confidence
184o3-mini1.2%estimated ± 3.6 pp, low confidence
185Ultravox v0.6 Llama 3.3 70B1.2%estimated ± 3.6 pp, low confidence
186GPT-4o1.2%estimated ± 3.6 pp, low confidence
187Mistral Large 31.2%estimated ± 3.6 pp, low confidence
188Mistral Medium 31.2%estimated ± 3.6 pp, low confidence
189DeepSeek V31.2%estimated ± 3.6 pp, low confidence
190Qwen3-Omni-30B-A3B-Thinking1.2%estimated ± 3.6 pp, low confidence
191Claude 3 Haiku1.2%estimated ± 3.6 pp, low confidence
192Gemma 4 E2B1.2%estimated ± 3.6 pp, low confidence
193Gemma 4 E4B1.2%estimated ± 3.6 pp, low confidence
194Exaone 4.0 1.2B1.2%estimated ± 3.6 pp, low confidence
195Granite-4.0-H-1B1.2%estimated ± 3.6 pp, low confidence
196Llama 3.1 405B1.2%estimated ± 3.6 pp, low confidence
197Llama 4 Maverick1.2%estimated ± 3.6 pp, low confidence
198GPT-4.1 nano1.2%estimated ± 3.6 pp, low confidence
199Qwen3-Omni-30B-A3B-Instruct1.2%estimated ± 3.6 pp, low confidence
200Llama 4 Scout1.2%estimated ± 3.6 pp, low confidence
201Gemini 2.5 Flash1.2%estimated ± 3.6 pp, low confidence
202Granite-4.0-H-350M1.2%estimated ± 3.6 pp, low confidence
203Nova Pro1.2%estimated ± 3.6 pp, low confidence
204Granite-4.0-350M1.2%estimated ± 3.6 pp, low confidence
205Nemotron Ultra 253B1.2%estimated ± 3.6 pp, low confidence
206Gemma 3 27B1.2%estimated ± 3.6 pp, low confidence
207Exaone 4.0 32B1.2%estimated ± 3.6 pp, low confidence
208LFM2.5-VL-1.6B-Extract1.2%estimated ± 3.6 pp, low confidence
209Phi-41.2%estimated ± 3.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General