benchgap
Calibration

Claw-Eval → BrowseComp

BrowseComp is estimated from Claw-Eval with a Michaelis–Menten curve fitted on 12 models measured on both: y = 2.0000·x / (1.11765 + x), R² = 0.33, cross-validated error 9.4 pp. It is used for 8 estimates.

Estimated modelClaw-EvalBrowseCompSource
GLM-5-Turbo55.8%66.6%estimated ± 9.4 pp, low confidence
K-EXAONE 2.077.7%82.0%estimated ± 9.4 pp, low confidence
LFM2.5-2.6B62.8%72.0%estimated ± 9.4 pp, low confidence
LLaDA2.2-mini57.2%67.7%estimated ± 9.4 pp, low confidence
MiMo-V2-Omni45.2%57.6%estimated ± 9.4 pp, low confidence
Ornith-1.0-35B69.8%76.9%estimated ± 9.4 pp, low confidence
Ornith-1.0-397B77.1%81.6%estimated ± 9.4 pp, low confidence
Ornith-1.0-9B63.1%72.2%estimated ± 9.4 pp, low confidence