benchgap
Agentic · tools

EnterpriseOps-Gym leaderboard

As of 2026-10-07, the highest measured score on EnterpriseOps-Gym is 51.1% by Claude Fable 5 (with fallback). 29 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: artificial-analysis.

#ModelScoreSource
1Gemini 3.7 Flash (high)56.7%estimated ± 1.8 pp, low confidence
2Claude Fable 5.1 (max with fallback)55.4%estimated ± 1.8 pp, low confidence
3Claude Sonnet 5.5 (max with fallback)55.4%estimated ± 1.8 pp, low confidence
4Claude Opus 5.5 (max with fallback)54.8%estimated ± 1.8 pp, low confidence
5Claude Opus 5 (max)53.6%estimated ± 1.8 pp, low confidence
6GPT-6 Astra (max)52.2%estimated ± 1.8 pp, low confidence
7GPT-6.1 Sol (max)51.6%estimated ± 1.8 pp, low confidence
8GPT-5.5 (xhigh)51.6%estimated ± 1.8 pp, low confidence
9Claude Fable 5 (with fallback)51.1%measured
10Gemini 3.7 Flash (medium)50.4%measured
11GPT-5.6 Sol (max)50.4%estimated ± 1.8 pp, medium confidence
12Gemini 3.5 Flash (high)50.1%measured
13Muse Spark 1.3 (xhigh)50.0%estimated ± 4.4 pp, low confidence
14DeepSeek V4 Pro 0813 (max)49.6%measured
15Grok 4.6 (high)48.3%measured
16Qwen3.8 Max (0902)47.6%measured
17Step 5 Preview47.2%measured
18Claude Fable 5.1 (xhigh with fallback)46.0%estimated ± 4.4 pp, medium confidence
19Claude Fable 5.1 (high with fallback)45.5%estimated ± 4.4 pp, medium confidence
20Kimi K3 (max)45.3%measured
21Qwen3.8 27B (xhigh)44.2%measured
22Qwen3.8 Max42.9%estimated ± 5.3 pp, low confidence
23Muse Spark 1.3 (max)42.8%estimated ± 5.3 pp, medium confidence
24Qwen3.8 2.4T A95B42.6%estimated ± 5.3 pp, medium confidence
25Gemini 4 Argon (high)42.5%estimated ± 6.2 pp, low confidence
26DeepSeek V4.1 Flash (max)42.4%estimated ± 6.2 pp, low confidence
27GPT-6 Astra (xhigh)42.4%estimated ± 6.2 pp, low confidence
28Grok 4.6 (xhigh)42.4%estimated ± 6.2 pp, low confidence
29GPT-6 Astra (high)42.3%estimated ± 6.2 pp, low confidence
30GPT-6.1 Sol (xhigh)42.3%estimated ± 6.2 pp, low confidence
31Grok 4.7 (xhigh)42.3%estimated ± 6.2 pp, low confidence
32Qwen3.8 27B (medium)42.3%estimated ± 5.3 pp, medium confidence
33GPT-6 Sol (max)42.3%estimated ± 6.2 pp, low confidence
34Mistral Large 4 Preview42.3%estimated ± 6.2 pp, low confidence
35MiMo-V2.6-Pro42.2%estimated ± 6.2 pp, low confidence
36GPT-6 Luna (max)42.1%estimated ± 6.2 pp, low confidence
37Gemini 3.8 Flash (high)41.9%estimated ± 5.3 pp, medium confidence
38K2 Horizon 375B A23B39.5%estimated ± 5.3 pp, medium confidence
39Inkling (xhigh)38.0%measured
40GLM-5.3 (max)36.4%measured
41Muse Glimmer (high)34.7%measured
42GLM-5.3-Flash33.2%measured
43MiniMax-M332.1%measured
44Nemotron 3 Ultra28.9%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Vision & documents