benchgap
Calibration

CyberGym → ApprenticeBench

ApprenticeBench is estimated from CyberGym with a offset logistic curve fitted on 7 models measured on both: y = 0.0341 + (0.3276 − 0.0341) / (1 + exp(−44.07·(x − 0.8179))), R² = 0.97, cross-validated error 4.0 pp. It is used for 18 estimates.

Estimated modelCyberGymApprenticeBenchSource
Atria Dawn Preview86.5%29.5%estimated ± 4.0 pp, low confidence
Claude Mythos 583.8%24.2%estimated ± 4.0 pp, medium confidence
Claude Mythos Preview83.1%22.2%estimated ± 4.0 pp, medium confidence
Claude Opus 4.7 (Adaptive)73.1%4.0%estimated ± 4.0 pp, medium confidence
DeepSeek V4.1 Flash88.1%31.0%estimated ± 4.0 pp, low confidence
DeepSeek V4 Flash 073176.7%6.2%estimated ± 4.0 pp, medium confidence
DeepSeek V4 Pro 081383.3%22.8%estimated ± 4.0 pp, medium confidence
Gemini 3.5 Flash Cyber83.2%22.5%estimated ± 4.0 pp, medium confidence
Gemini 3.8 Flash Cyber86.2%29.1%estimated ± 4.0 pp, low confidence
GLM-5.384.5%25.9%estimated ± 4.0 pp, medium confidence
Hy4 preview78.4%8.8%estimated ± 4.0 pp, medium confidence
Ling 3.1 Flash87.9%30.9%estimated ± 4.0 pp, low confidence
MiMo-V2.6-Flash95.1%32.7%estimated ± 4.0 pp, low confidence
MiMo-V2.6-Pro94.0%32.6%estimated ± 4.0 pp, low confidence
Muse Spark43.5%3.4%estimated ± 4.0 pp, low confidence
Muse Spark 1.159.0%3.4%estimated ± 4.0 pp, low confidence
Fugu Cyber86.9%30.0%estimated ± 4.0 pp, low confidence
Step 5 Preview84.7%26.4%estimated ± 4.0 pp, low confidence