benchgap
Calibration

OSWorld-Verified → OSWorld 2.0

OSWorld 2.0 is estimated from OSWorld-Verified with a linear curve fitted on 9 models measured on both: y = 1.1206·x + -0.7486, R² = 0.80, cross-validated error 3.6 pp. It is used for 25 estimates.

Estimated modelOSWorld-VerifiedOSWorld 2.0Source
Claude Fable 585.0%20.4%estimated ± 3.6 pp, high confidence
Claude Mythos 585.0%20.4%estimated ± 3.6 pp, high confidence
Claude Opus 4.566.3%0.0%estimated ± 3.6 pp, medium confidence
Claude Opus 4.672.7%6.6%estimated ± 3.6 pp, high confidence
Claude Sonnet 4.561.4%0.0%estimated ± 3.6 pp, medium confidence
Claude Sonnet 581.2%16.1%estimated ± 3.6 pp, high confidence
Gemini 3.5 Flash78.4%13.0%estimated ± 3.6 pp, high confidence
Gemini 3.5 Flash-Lite74.0%8.1%estimated ± 3.6 pp, high confidence
Gemini 3.6 Flash83.0%18.1%estimated ± 3.6 pp, high confidence
GPT-5.247.3%0.0%estimated ± 3.6 pp, medium confidence
GPT-5.3 Codex64.7%0.0%estimated ± 3.6 pp, medium confidence
GPT-5.475.0%9.2%estimated ± 3.6 pp, high confidence
GPT-5.4 mini72.1%5.9%estimated ± 3.6 pp, high confidence
GPT-5.4 nano39.0%0.0%estimated ± 3.6 pp, medium confidence
Holo3-122B-A10B78.9%13.5%estimated ± 3.6 pp, high confidence
Holo3-35B-A3B82.6%17.7%estimated ± 3.6 pp, high confidence
MiMo-V2.6-Flash80.8%15.7%estimated ± 3.6 pp, high confidence
MiMo-V2.6-Pro82.0%17.0%estimated ± 3.6 pp, high confidence
Muse Glimmer 30B65.9%0.0%estimated ± 3.6 pp, medium confidence
Qwen3.5-122B-A10B58.0%0.0%estimated ± 3.6 pp, medium confidence
Qwen3.5-27B56.2%0.0%estimated ± 3.6 pp, medium confidence
Qwen3.5-35B-A3B54.5%0.0%estimated ± 3.6 pp, medium confidence
Qwen3.8-27B84.3%19.6%estimated ± 3.6 pp, high confidence
UI-Mate-27B77.0%11.4%estimated ± 3.6 pp, high confidence
UI-Mate-9B66.2%0.0%estimated ± 3.6 pp, medium confidence