benchgap
Calibration

BrowseComp → Toolathlon

Toolathlon is estimated from BrowseComp with a Hill curve fitted on 12 models measured on both: y = 0.0000 + (0.6180 − 0.0000)·x^6.00 / (0.60424^6.00 + x^6.00), R² = 0.91, cross-validated error 3.5 pp. It is used for 28 estimates.

Estimated modelBrowseCompToolathlonSource
Agents-A175.5%48.9%estimated ± 3.5 pp, high confidence
Agents-A1-4B66.8%39.9%estimated ± 3.5 pp, high confidence
Atria Dawn Preview92.5%57.3%estimated ± 3.5 pp, medium confidence
Claude Mythos 588.0%55.9%estimated ± 3.5 pp, high confidence
Claude Opus 4.7 (Adaptive)79.3%51.7%estimated ± 3.5 pp, high confidence
dots3-note Preview83.3%53.9%estimated ± 3.5 pp, high confidence
GLM-4.752.0%17.9%estimated ± 3.5 pp, medium confidence
GLM-5.168.0%41.4%estimated ± 3.5 pp, high confidence
GPT-5.265.8%38.6%estimated ± 3.5 pp, high confidence
GPT-5.4 Pro89.3%56.4%estimated ± 3.5 pp, high confidence
GPT-5.5 Pro90.1%56.6%estimated ± 3.5 pp, high confidence
Inkling77.1%50.2%estimated ± 3.5 pp, high confidence
Inkling-Small77.4%50.4%estimated ± 3.5 pp, high confidence
Kimi K2.5 (Reasoning)60.6%31.2%estimated ± 3.5 pp, high confidence
Ling 3.0 Flash72.2%46.0%estimated ± 3.5 pp, high confidence
LongCat-Flash-Lite-Sparse48.6%13.2%estimated ± 3.5 pp, medium confidence
MiniMax M383.5%54.0%estimated ± 3.5 pp, high confidence
Nemotron 3.5 Lightning 30B A3B NVFP436.8%3.0%estimated ± 3.5 pp, medium confidence
Nemotron 3 Ultra44.4%8.4%estimated ± 3.5 pp, medium confidence
Ornith-1.5-35B-A3B67.6%40.9%estimated ± 3.5 pp, high confidence
Ornith-1.5-397B86.6%55.4%estimated ± 3.5 pp, high confidence
Ornith-1.5-9B56.4%24.6%estimated ± 3.5 pp, medium confidence
Qwen3.5-122B-A10B63.8%35.9%estimated ± 3.5 pp, high confidence
Qwen3.5-27B61.0%31.8%estimated ± 3.5 pp, high confidence
Qwen3.5-35B-A3B61.0%31.8%estimated ± 3.5 pp, high confidence
Beam77.4%50.4%estimated ± 3.5 pp, high confidence
Solar Pro 449.2%13.9%estimated ± 3.5 pp, medium confidence
Step 5 Preview88.7%56.2%estimated ± 3.5 pp, high confidence