benchgap
Calibration

ExploitGym → BrowseComp

BrowseComp is estimated from ExploitGym with a linear curve fitted on 6 models measured on both: y = 0.2866·x + 0.8067, R² = 0.93, cross-validated error 1.8 pp. It is used for 9 estimates.

Estimated modelExploitGymBrowseCompSource
Claude Mythos Preview17.5%85.7%estimated ± 1.8 pp, medium confidence
DeepSeek V4.1 Flash15.3%85.1%estimated ± 1.8 pp, medium confidence
GLM-5.315.0%85.0%estimated ± 1.8 pp, medium confidence
GPT-6.1 Sol35.1%90.7%estimated ± 1.8 pp, medium confidence
GPT-6 Luna11.6%84.0%estimated ± 1.8 pp, medium confidence
GPT-6 Sol22.1%87.0%estimated ± 1.8 pp, medium confidence
MiMo-V2.6-Flash6.0%82.4%estimated ± 1.8 pp, medium confidence
MiMo-V2.6-Pro17.8%85.8%estimated ± 1.8 pp, medium confidence
Muse Spark 1.10.8%80.9%estimated ± 1.8 pp, low confidence