Calibration
AA AutomationBench → OSWorld 2.0
OSWorld 2.0 is estimated from AA AutomationBench with a Michaelis–Menten curve fitted on 8 models measured on both: y = 2.0000·x / (1.57019 + x), R² = 0.69, cross-validated error 13.3 pp. It is used for 9 estimates.
| Estimated model | AA AutomationBench | OSWorld 2.0 | Source |
|---|---|---|---|
| Claude Haiku 5.5 | 35.4% | 36.8% | estimated ± 13.3 pp, low confidence |
| DeepSeek V4.1 Flash | 68.9% | 61.0% | estimated ± 13.3 pp, low confidence |
| GPT-6.1 Sol | 64.9% | 58.5% | estimated ± 13.3 pp, low confidence |
| GPT-6 Luna | 53.2% | 50.6% | estimated ± 13.3 pp, low confidence |
| Grok 4.6 | 66.7% | 59.6% | estimated ± 13.3 pp, low confidence |
| Grok 4.7 | 65.6% | 58.9% | estimated ± 13.3 pp, low confidence |
| Inkling | 5.0% | 6.2% | estimated ± 13.3 pp, low confidence |
| Mistral Large 4 | 59.9% | 55.2% | estimated ± 13.3 pp, low confidence |
| Nemotron 3 Ultra | 3.0% | 3.7% | estimated ± 13.3 pp, low confidence |