benchgap
Calibration

OSWorld-Verified → DeepSearchQA

DeepSearchQA is estimated from OSWorld-Verified with a Michaelis–Menten curve fitted on 6 models measured on both: y = 2.0000·x / (1.07697 + x), R² = 0.33, cross-validated error 8.5 pp. It is used for 8 estimates.

Estimated modelOSWorld-VerifiedDeepSearchQASource
Claude Opus 4.566.3%76.2%estimated ± 8.5 pp, low confidence
Claude Sonnet 4.561.4%72.6%estimated ± 8.5 pp, low confidence
GPT-5.3 Codex64.7%75.1%estimated ± 8.5 pp, low confidence
Holo3-122B-A10B78.9%84.5%estimated ± 8.5 pp, low confidence
Holo3-35B-A3B82.6%86.8%estimated ± 8.5 pp, low confidence
Qwen3.8 Max86.1%88.9%estimated ± 8.5 pp, low confidence
UI-Mate-27B77.0%83.4%estimated ± 8.5 pp, low confidence
UI-Mate-9B66.2%76.1%estimated ± 8.5 pp, low confidence