Calibration
OSWorld-Verified → DeepSearchQA
DeepSearchQA is estimated from OSWorld-Verified with a Michaelis–Menten curve fitted on 6 models measured on both: y = 2.0000·x / (1.07697 + x), R² = 0.33, cross-validated error 8.5 pp. It is used for 8 estimates.
| Estimated model | OSWorld-Verified | DeepSearchQA | Source |
|---|---|---|---|
| Claude Opus 4.5 | 66.3% | 76.2% | estimated ± 8.5 pp, low confidence |
| Claude Sonnet 4.5 | 61.4% | 72.6% | estimated ± 8.5 pp, low confidence |
| GPT-5.3 Codex | 64.7% | 75.1% | estimated ± 8.5 pp, low confidence |
| Holo3-122B-A10B | 78.9% | 84.5% | estimated ± 8.5 pp, low confidence |
| Holo3-35B-A3B | 82.6% | 86.8% | estimated ± 8.5 pp, low confidence |
| Qwen3.8 Max | 86.1% | 88.9% | estimated ± 8.5 pp, low confidence |
| UI-Mate-27B | 77.0% | 83.4% | estimated ± 8.5 pp, low confidence |
| UI-Mate-9B | 66.2% | 76.1% | estimated ± 8.5 pp, low confidence |