benchgap
Calibration

AA AutomationBench → OSWorld 2.0

OSWorld 2.0 is estimated from AA AutomationBench with a Michaelis–Menten curve fitted on 8 models measured on both: y = 2.0000·x / (1.57019 + x), R² = 0.69, cross-validated error 13.3 pp. It is used for 9 estimates.

Estimated modelAA AutomationBenchOSWorld 2.0Source
Claude Haiku 5.535.4%36.8%estimated ± 13.3 pp, low confidence
DeepSeek V4.1 Flash68.9%61.0%estimated ± 13.3 pp, low confidence
GPT-6.1 Sol64.9%58.5%estimated ± 13.3 pp, low confidence
GPT-6 Luna53.2%50.6%estimated ± 13.3 pp, low confidence
Grok 4.666.7%59.6%estimated ± 13.3 pp, low confidence
Grok 4.765.6%58.9%estimated ± 13.3 pp, low confidence
Inkling5.0%6.2%estimated ± 13.3 pp, low confidence
Mistral Large 459.9%55.2%estimated ± 13.3 pp, low confidence
Nemotron 3 Ultra3.0%3.7%estimated ± 13.3 pp, low confidence