benchgap
Calibration

CyberGym → AutomationBench

AutomationBench is estimated from CyberGym with a offset logistic curve fitted on 10 models measured on both: y = 0.2844 + (0.5326 − 0.2844) / (1 + exp(−198.17·(x − 0.8413))), R² = 0.96, cross-validated error 5.8 pp. It is used for 18 estimates.

Estimated modelCyberGymAutomationBenchSource
Claude Mythos 583.8%36.9%estimated ± 5.8 pp, medium confidence
Claude Mythos Preview83.1%31.3%estimated ± 5.8 pp, medium confidence
Claude Opus 4.550.6%28.4%estimated ± 5.8 pp, low confidence
Claude Opus 4.666.6%28.4%estimated ± 5.8 pp, low confidence
Claude Opus 4.7 (Adaptive)73.1%28.4%estimated ± 5.8 pp, low confidence
Claude Sonnet 4.665.2%28.4%estimated ± 5.8 pp, low confidence
Gemini 3.5 Flash Cyber83.2%31.8%estimated ± 5.8 pp, medium confidence
Gemini 3.8 Flash Cyber86.2%52.9%estimated ± 5.8 pp, medium confidence
GLM-543.2%28.4%estimated ± 5.8 pp, low confidence
GLM-5.168.7%28.4%estimated ± 5.8 pp, low confidence
GPT-5.479.0%28.4%estimated ± 5.8 pp, medium confidence
GPT-5.581.8%28.7%estimated ± 5.8 pp, medium confidence
GPT-5.6 Luna77.9%28.4%estimated ± 5.8 pp, medium confidence
GPT-5.6 Sol84.5%45.2%estimated ± 5.8 pp, medium confidence
GPT-5.6 Terra81.8%28.7%estimated ± 5.8 pp, medium confidence
Muse Spark43.5%28.4%estimated ± 5.8 pp, low confidence
Muse Spark 1.159.0%28.4%estimated ± 5.8 pp, low confidence
Fugu Cyber86.9%53.2%estimated ± 5.8 pp, medium confidence