benchgap
Calibration

BrowseComp → AutomationBench

AutomationBench is estimated from BrowseComp with a inverse Michaelis–Menten curve fitted on 7 models measured on both: y = 0.24873·(x − 0.0000) / (0.0000 + 1.4849 − x), R² = 0.31, cross-validated error 10.3 pp. It is used for 14 estimates.

Estimated modelBrowseCompAutomationBenchSource
Agents-A175.5%25.7%estimated ± 10.3 pp, low confidence
Agents-A1-4B66.8%20.3%estimated ± 10.3 pp, low confidence
dots3-note Preview83.3%31.8%estimated ± 10.3 pp, low confidence
GPT-5.265.8%19.8%estimated ± 10.3 pp, low confidence
GPT-5.4 Pro89.3%37.5%estimated ± 10.3 pp, low confidence
GPT-5.5 Pro90.1%38.4%estimated ± 10.3 pp, low confidence
LongCat-Flash-Lite-Sparse48.6%12.1%estimated ± 10.3 pp, low confidence
Ornith-1.5-35B-A3B67.6%20.8%estimated ± 10.3 pp, low confidence
Ornith-1.5-397B86.6%34.8%estimated ± 10.3 pp, low confidence
Ornith-1.5-9B56.4%15.2%estimated ± 10.3 pp, low confidence
Qwen3.5-27B61.0%17.3%estimated ± 10.3 pp, low confidence
Qwen3.5-35B-A3B61.0%17.3%estimated ± 10.3 pp, low confidence
Qwen3.5 397B62.0%17.8%estimated ± 10.3 pp, low confidence
Beam77.4%27.1%estimated ± 10.3 pp, low confidence