benchgap
Calibration

BrowseComp → ResearchClawBench

ResearchClawBench is estimated from BrowseComp with a Michaelis–Menten + offset curve fitted on 9 models measured on both: y = 0.0000 + 0.6410·x / (2.03915 + x), R² = 0.47, cross-validated error 2.2 pp. It is used for 11 estimates.

Estimated modelBrowseCompResearchClawBenchSource
Agents-A175.5%17.3%estimated ± 2.2 pp, medium confidence
Agents-A1-4B66.8%15.8%estimated ± 2.2 pp, medium confidence
Atria Dawn Preview92.5%20.0%estimated ± 2.2 pp, low confidence
Claude Mythos 588.0%19.3%estimated ± 2.2 pp, low confidence
GLM-4.752.0%13.0%estimated ± 2.2 pp, low confidence
GPT-5.265.8%15.6%estimated ± 2.2 pp, medium confidence
GPT-5.4 Pro89.3%19.5%estimated ± 2.2 pp, low confidence
GPT-5.5 Pro90.1%19.6%estimated ± 2.2 pp, low confidence
Kimi K2.5 (Reasoning)60.6%14.7%estimated ± 2.2 pp, medium confidence
Qwen3.5-27B61.0%14.8%estimated ± 2.2 pp, medium confidence
Qwen3.5-35B-A3B61.0%14.8%estimated ± 2.2 pp, medium confidence