benchgap
Calibration

Gert Labs → OSWorld 2.0

OSWorld 2.0 is estimated from Gert Labs with a linear curve fitted on 5 models measured on both: y = 0.7639·x + -0.3853, R² = 0.76, cross-validated error 4.5 pp. It is used for 38 estimates.

Estimated modelGert LabsOSWorld 2.0Source
Claude 4 Sonnet39.7%0.0%estimated ± 4.5 pp, low confidence
DeepSeek V3.229.6%0.0%estimated ± 4.5 pp, low confidence
Gemini 2.5 Pro42.0%0.0%estimated ± 4.5 pp, low confidence
Gemini 3.1 Flash-Lite38.5%0.0%estimated ± 4.5 pp, low confidence
Gemini 3.1 Pro56.9%4.9%estimated ± 4.5 pp, medium confidence
Gemini 3 Flash56.6%4.7%estimated ± 4.5 pp, low confidence
Gemini 3 Pro63.2%9.8%estimated ± 4.5 pp, medium confidence
Gemma 4 31B35.3%0.0%estimated ± 4.5 pp, low confidence
GLM-4.740.0%0.0%estimated ± 4.5 pp, low confidence
GLM-551.0%0.4%estimated ± 4.5 pp, low confidence
GLM-5.160.1%7.4%estimated ± 4.5 pp, medium confidence
GLM-5V-Turbo30.8%0.0%estimated ± 4.5 pp, low confidence
GPT-4.125.7%0.0%estimated ± 4.5 pp, low confidence
GPT-5.141.2%0.0%estimated ± 4.5 pp, low confidence
GPT-5.1-Codex49.7%0.0%estimated ± 4.5 pp, low confidence
GPT-5.2-Codex51.8%1.0%estimated ± 4.5 pp, low confidence
GPT-OSS 120B29.6%0.0%estimated ± 4.5 pp, low confidence
Grok 442.3%0.0%estimated ± 4.5 pp, low confidence
Grok 4.1 Fast47.3%0.0%estimated ± 4.5 pp, low confidence
Grok 4.2038.4%0.0%estimated ± 4.5 pp, low confidence
Grok 4.343.9%0.0%estimated ± 4.5 pp, low confidence
Grok Build 0.149.2%0.0%estimated ± 4.5 pp, low confidence
Hy3 Preview36.9%0.0%estimated ± 4.5 pp, low confidence
Kimi K2.545.9%0.0%estimated ± 4.5 pp, low confidence
Kimi K2.5 (Reasoning)32.6%0.0%estimated ± 4.5 pp, low confidence
MiMo-V2.546.9%0.0%estimated ± 4.5 pp, low confidence
MiMo-V2.5-Pro62.7%9.4%estimated ± 4.5 pp, medium confidence
MiMo-V2-Pro36.7%0.0%estimated ± 4.5 pp, low confidence
MiniMax M2.740.4%0.0%estimated ± 4.5 pp, low confidence
Mistral Medium 3.5 128B39.1%0.0%estimated ± 4.5 pp, low confidence
Qwen3.5 397B46.8%0.0%estimated ± 4.5 pp, low confidence
Qwen3.6-27B54.8%3.4%estimated ± 4.5 pp, low confidence
Qwen3.6-35B-A3B42.7%0.0%estimated ± 4.5 pp, low confidence
Qwen3.6 Plus50.6%0.1%estimated ± 4.5 pp, low confidence
Qwen3.7 Max64.3%10.6%estimated ± 4.5 pp, medium confidence
Qwen3 Max43.7%0.0%estimated ± 4.5 pp, low confidence
Step 3.7 Flash51.6%0.9%estimated ± 4.5 pp, low confidence
Trinity-Large-Thinking32.6%0.0%estimated ± 4.5 pp, low confidence