benchgap
Calibration

ExploitGym → AA ITBench

AA ITBench is estimated from ExploitGym with a offset logistic curve fitted on 7 models measured on both: y = 0.4626 + (0.5240 − 0.4626) / (1 + exp(−105.19·(x − 0.2205))), R² = 0.63, cross-validated error 3.9 pp. It is used for 6 estimates.

Estimated modelExploitGymAA ITBenchSource
Claude Mythos Preview17.5%46.3%estimated ± 3.9 pp, medium confidence
GPT-5.6 Luna12.4%46.3%estimated ± 3.9 pp, low confidence
GPT-6.1 Sol35.1%52.4%estimated ± 3.9 pp, medium confidence
GPT-6 Luna11.6%46.3%estimated ± 3.9 pp, low confidence
MiMo-V2.6-Flash6.0%46.3%estimated ± 3.9 pp, low confidence
MiMo-V2.6-Pro17.8%46.3%estimated ± 3.9 pp, medium confidence