benchgap
Calibration

BrowseComp → Claw-Eval

Claw-Eval is estimated from BrowseComp with a offset logistic curve fitted on 12 models measured on both: y = 0.6240 + (0.8170 − 0.6240) / (1 + exp(−200.00·(x − 0.8368))), R² = 0.51, cross-validated error 7.9 pp. It is used for 11 estimates.

Estimated modelBrowseCompClaw-EvalSource
Claude Mythos 588.0%81.7%estimated ± 7.9 pp, low confidence
Claude Sonnet 584.7%79.5%estimated ± 7.9 pp, medium confidence
GPT-5.4 Pro89.3%81.7%estimated ± 7.9 pp, low confidence
GPT-5.5 Pro90.1%81.7%estimated ± 7.9 pp, low confidence
GPT-5.6 Luna83.3%68.5%estimated ± 7.9 pp, medium confidence
GPT-5.6 Sol92.2%81.7%estimated ± 7.9 pp, low confidence
GPT-5.6 Terra87.5%81.7%estimated ± 7.9 pp, low confidence
GPT-6 Astra91.5%81.7%estimated ± 7.9 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP436.8%62.4%estimated ± 7.9 pp, low confidence
Nemotron 3 Ultra44.4%62.4%estimated ± 7.9 pp, low confidence
Qwen3.5-122B-A10B63.8%62.4%estimated ± 7.9 pp, medium confidence