benchgap
Calibration

BrowseComp → HLE w/ tools

HLE w/ tools is estimated from BrowseComp with a offset logistic curve fitted on 12 models measured on both: y = 0.3318 + (0.5959 − 0.3318) / (1 + exp(−26.96·(x − 0.7512))), R² = 0.92, cross-validated error 4.6 pp. It is used for 28 estimates.

Estimated modelBrowseCompHLE w/ toolsSource
Agents-A1-4B66.8%35.7%estimated ± 4.6 pp, high confidence
Claude Mythos 588.0%58.8%estimated ± 4.6 pp, high confidence
Claude Opus 4.683.7%57.2%estimated ± 4.6 pp, high confidence
Claude Opus 4.7 (Adaptive)79.3%53.1%estimated ± 4.6 pp, high confidence
Claude Opus 4.884.3%57.5%estimated ± 4.6 pp, high confidence
GLM-4.752.0%33.2%estimated ± 4.6 pp, high confidence
GLM-5.168.0%36.6%estimated ± 4.6 pp, high confidence
GPT-5.265.8%35.2%estimated ± 4.6 pp, high confidence
GPT-5.482.7%56.6%estimated ± 4.6 pp, high confidence
GPT-5.4 Pro89.3%59.0%estimated ± 4.6 pp, high confidence
GPT-5.584.4%57.6%estimated ± 4.6 pp, high confidence
GPT-5.5 Pro90.1%59.1%estimated ± 4.6 pp, high confidence
GPT-5.6 Luna83.3%57.0%estimated ± 4.6 pp, high confidence
GPT-5.6 Sol92.2%59.3%estimated ± 4.6 pp, medium confidence
GPT-5.6 Terra87.5%58.7%estimated ± 4.6 pp, high confidence
Inkling77.1%49.8%estimated ± 4.6 pp, high confidence
Inkling-Small77.4%50.3%estimated ± 4.6 pp, high confidence
Kimi K2.5 (Reasoning)60.6%33.7%estimated ± 4.6 pp, high confidence
Kimi K391.2%59.3%estimated ± 4.6 pp, high confidence
LongCat-Flash-Lite-Sparse48.6%33.2%estimated ± 4.6 pp, high confidence
MiniMax M383.5%57.1%estimated ± 4.6 pp, high confidence
Nemotron 3.5 Lightning 30B A3B NVFP436.8%33.2%estimated ± 4.6 pp, medium confidence
Qwen3.5-122B-A10B63.8%34.4%estimated ± 4.6 pp, high confidence
Qwen3.5-27B61.0%33.8%estimated ± 4.6 pp, high confidence
Qwen3.5-35B-A3B61.0%33.8%estimated ± 4.6 pp, high confidence
Beam77.4%50.3%estimated ± 4.6 pp, high confidence
Solar Pro 449.2%33.2%estimated ± 4.6 pp, high confidence
Step 5 Preview88.7%58.9%estimated ± 4.6 pp, high confidence