benchgap
Calibration

OSWorld-Verified → BrowseComp

BrowseComp is estimated from OSWorld-Verified with a offset logistic curve fitted on 13 models measured on both: y = 0.6282 + (0.8382 − 0.6282) / (1 + exp(−70.17·(x − 0.6443))), R² = 0.95, cross-validated error 3.1 pp. It is used for 14 estimates.

Estimated modelOSWorld-VerifiedBrowseCompSource
Claude Opus 4.566.3%79.4%estimated ± 3.1 pp, high confidence
Claude Sonnet 4.561.4%65.1%estimated ± 3.1 pp, high confidence
Gemini 3.5 Flash78.4%83.8%estimated ± 3.1 pp, high confidence
Gemini 3.5 Flash-Lite74.0%83.8%estimated ± 3.1 pp, high confidence
Gemini 3.6 Flash83.0%83.8%estimated ± 3.1 pp, high confidence
GPT-5.3 Codex64.7%74.3%estimated ± 3.1 pp, high confidence
GPT-5.4 mini72.1%83.7%estimated ± 3.1 pp, high confidence
GPT-5.4 nano39.0%62.8%estimated ± 3.1 pp, medium confidence
Holo3-122B-A10B78.9%83.8%estimated ± 3.1 pp, high confidence
Holo3-35B-A3B82.6%83.8%estimated ± 3.1 pp, high confidence
Muse Glimmer 30B65.9%78.3%estimated ± 3.1 pp, high confidence
Qwen3.8-27B84.3%83.8%estimated ± 3.1 pp, high confidence
UI-Mate-27B77.0%83.8%estimated ± 3.1 pp, high confidence
UI-Mate-9B66.2%79.1%estimated ± 3.1 pp, high confidence