benchgap
Calibration

BrowseComp → WideResearch

WideResearch is estimated from BrowseComp with a Michaelis–Menten curve fitted on 9 models measured on both: y = 1.4235·x / (0.66355 + x), R² = 0.75, cross-validated error 4.6 pp. It is used for 20 estimates.

Estimated modelBrowseCompWideResearchSource
Agents-A1-4B66.8%71.4%estimated ± 4.6 pp, high confidence
Claude Mythos 588.0%81.2%estimated ± 4.6 pp, high confidence
Claude Opus 4.884.3%79.7%estimated ± 4.6 pp, high confidence
GLM-4.752.0%62.5%estimated ± 4.6 pp, medium confidence
GLM-5.168.0%72.0%estimated ± 4.6 pp, high confidence
GPT-5.4 Pro89.3%81.7%estimated ± 4.6 pp, high confidence
GPT-5.5 Pro90.1%82.0%estimated ± 4.6 pp, high confidence
GPT-5.6 Luna83.3%79.2%estimated ± 4.6 pp, high confidence
GPT-5.6 Sol92.2%82.8%estimated ± 4.6 pp, high confidence
GPT-5.6 Terra87.5%81.0%estimated ± 4.6 pp, high confidence
Inkling77.1%76.5%estimated ± 4.6 pp, high confidence
Kimi K2.5 (Reasoning)60.6%67.9%estimated ± 4.6 pp, high confidence
LongCat-Flash-Lite-Sparse48.6%60.2%estimated ± 4.6 pp, medium confidence
MiniMax M383.5%79.3%estimated ± 4.6 pp, high confidence
Nemotron 3.5 Lightning 30B A3B NVFP436.8%50.8%estimated ± 4.6 pp, medium confidence
Qwen3.5-122B-A10B63.8%69.8%estimated ± 4.6 pp, high confidence
Qwen3.5-27B61.0%68.2%estimated ± 4.6 pp, high confidence
Qwen3.5-35B-A3B61.0%68.2%estimated ± 4.6 pp, high confidence
Beam77.4%76.6%estimated ± 4.6 pp, high confidence
Solar Pro 449.2%60.6%estimated ± 4.6 pp, medium confidence