benchgap
Calibration

AutomationBench → EnterpriseOps-Gym

EnterpriseOps-Gym is estimated from AutomationBench with a Michaelis–Menten curve fitted on 11 models measured on both: y = 0.4321·x / (0.01355 + x), R² = 0.37, cross-validated error 6.2 pp. It is used for 11 estimates.

Estimated modelAutomationBenchEnterpriseOps-GymSource
GPT-6 Astra (high)66.6%42.3%estimated ± 6.2 pp, low confidence
GPT-6 Astra (xhigh)67.2%42.4%estimated ± 6.2 pp, low confidence
Gemini 4 Argon (high)77.5%42.5%estimated ± 6.2 pp, low confidence
GPT-6 Sol (max)61.6%42.3%estimated ± 6.2 pp, low confidence
MiMo-V2.6-Pro58.6%42.2%estimated ± 6.2 pp, low confidence
DeepSeek V4.1 Flash (max)68.9%42.4%estimated ± 6.2 pp, low confidence
Mistral Large 4 Preview59.9%42.3%estimated ± 6.2 pp, low confidence
Grok 4.7 (xhigh)65.6%42.3%estimated ± 6.2 pp, low confidence
GPT-6 Luna (max)53.2%42.1%estimated ± 6.2 pp, low confidence
GPT-6.1 Sol (xhigh)66.6%42.3%estimated ± 6.2 pp, low confidence
Grok 4.6 (xhigh)67.0%42.4%estimated ± 6.2 pp, low confidence