benchgap
Calibration

CyberGym → Agents' Last Exam

Agents' Last Exam is estimated from CyberGym with a offset logistic curve fitted on 8 models measured on both: y = 0.2401 + (0.3035 − 0.2401) / (1 + exp(−182.53·(x − 0.8388))), R² = 0.80, cross-validated error 2.1 pp. It is used for 20 estimates.

Estimated modelCyberGymAgents' Last ExamSource
Atria Dawn Preview86.5%30.3%estimated ± 2.1 pp, high confidence
Claude Mythos 583.8%27.0%estimated ± 2.1 pp, high confidence
Claude Mythos Preview83.1%25.2%estimated ± 2.1 pp, high confidence
Claude Opus 4.550.6%24.0%estimated ± 2.1 pp, medium confidence
Claude Opus 4.666.6%24.0%estimated ± 2.1 pp, medium confidence
Claude Opus 4.7 (Adaptive)73.1%24.0%estimated ± 2.1 pp, medium confidence
Claude Sonnet 4.665.2%24.0%estimated ± 2.1 pp, medium confidence
Gemini 3.5 Flash Cyber83.2%25.4%estimated ± 2.1 pp, high confidence
Gemini 3.8 Flash Cyber86.2%30.3%estimated ± 2.1 pp, high confidence
GLM-543.2%24.0%estimated ± 2.1 pp, medium confidence
GLM-5.168.7%24.0%estimated ± 2.1 pp, medium confidence
GPT-5.479.0%24.0%estimated ± 2.1 pp, high confidence
GPT-5.581.8%24.2%estimated ± 2.1 pp, high confidence
GPT-5.6 Luna77.9%24.0%estimated ± 2.1 pp, high confidence
GPT-5.6 Sol84.5%28.8%estimated ± 2.1 pp, high confidence
GPT-5.6 Terra81.8%24.2%estimated ± 2.1 pp, high confidence
Ling 3.1 Flash87.9%30.3%estimated ± 2.1 pp, high confidence
Muse Spark43.5%24.0%estimated ± 2.1 pp, medium confidence
Muse Spark 1.159.0%24.0%estimated ± 2.1 pp, medium confidence
Fugu Cyber86.9%30.3%estimated ± 2.1 pp, high confidence