benchgap
Calibration

BrowseComp → DeepSearchQA

DeepSearchQA is estimated from BrowseComp with a Hill curve fitted on 11 models measured on both: y = 0.7463 + (1.2000 − 0.7463)·x^6.00 / (0.97379^6.00 + x^6.00), R² = 0.31, cross-validated error 8.1 pp. It is used for 17 estimates.

Estimated modelBrowseCompDeepSearchQASource
Agents-A175.5%82.7%estimated ± 8.1 pp, low confidence
Agents-A1-4B66.8%78.9%estimated ± 8.1 pp, low confidence
Claude Mythos 588.0%90.6%estimated ± 8.1 pp, low confidence
GLM-4.752.0%75.7%estimated ± 8.1 pp, low confidence
GPT-5.265.8%78.6%estimated ± 8.1 pp, low confidence
GPT-5.4 Pro89.3%91.6%estimated ± 8.1 pp, low confidence
GPT-5.5 Pro90.1%92.1%estimated ± 8.1 pp, low confidence
Kimi K2.5 (Reasoning)60.6%77.1%estimated ± 8.1 pp, low confidence
LongCat-Flash-Lite-Sparse48.6%75.3%estimated ± 8.1 pp, low confidence
Ornith-1.5-35B-A3B67.6%79.2%estimated ± 8.1 pp, low confidence
Ornith-1.5-397B86.6%89.6%estimated ± 8.1 pp, low confidence
Ornith-1.5-9B56.4%76.3%estimated ± 8.1 pp, low confidence
Qwen3.5-27B61.0%77.2%estimated ± 8.1 pp, low confidence
Qwen3.5-35B-A3B61.0%77.2%estimated ± 8.1 pp, low confidence
Qwen3.5 397B62.0%77.5%estimated ± 8.1 pp, low confidence
Solar Pro 449.2%75.4%estimated ± 8.1 pp, low confidence
Step 5 Preview88.7%91.1%estimated ± 8.1 pp, low confidence