Calibration
OSWorld 2.0 → JobBench
JobBench is estimated from OSWorld 2.0 with a Michaelis–Menten curve fitted on 7 models measured on both: y = 0.7273·x / (0.07457 + x), R² = 0.80, cross-validated error 4.9 pp. It is used for 13 estimates.
| Estimated model | OSWorld 2.0 | JobBench | Source |
|---|---|---|---|
| Claude Opus 4.7 | 13.9% | 47.3% | estimated ± 4.9 pp, medium confidence |
| Claude Opus 4.8 | 20.6% | 53.4% | estimated ± 4.9 pp, medium confidence |
| Gemini 3.7 Flash | 47.9% | 62.9% | estimated ± 4.9 pp, medium confidence |
| Gemini 3.8 Flash | 59.0% | 64.6% | estimated ± 4.9 pp, medium confidence |
| Gemini 4 Argon | 69.2% | 65.7% | estimated ± 4.9 pp, low confidence |
| GPT-5.6 Luna | 45.6% | 62.5% | estimated ± 4.9 pp, medium confidence |
| GPT-5.6 Sol | 62.6% | 65.0% | estimated ± 4.9 pp, medium confidence |
| GPT-5.6 Terra | 50.2% | 63.3% | estimated ± 4.9 pp, medium confidence |
| GPT-6 Astra | 72.6% | 66.0% | estimated ± 4.9 pp, low confidence |
| GPT-6 Sol | 60.5% | 64.8% | estimated ± 4.9 pp, medium confidence |
| Kimi K2.6 | 4.6% | 27.7% | estimated ± 4.9 pp, low confidence |
| MiniMax M3 | 4.6% | 27.7% | estimated ± 4.9 pp, low confidence |
| Qwen3.7 Plus | 2.8% | 19.9% | estimated ± 4.9 pp, low confidence |