Calibration
OSWorld 2.0 → Toolathlon-Verified
Toolathlon-Verified is estimated from OSWorld 2.0 with a Michaelis–Menten curve fitted on 5 models measured on both: y = 0.8288·x / (0.02655 + x), R² = 0.96, cross-validated error 1.0 pp. It is used for 18 estimates.
| Estimated model | OSWorld 2.0 | Toolathlon-Verified | Source |
|---|---|---|---|
| Claude Opus 4.7 | 13.9% | 69.6% | estimated ± 1.0 pp, low confidence |
| Claude Opus 4.7 (Adaptive) | 18.2% | 72.3% | estimated ± 1.0 pp, low confidence |
| Claude Opus 4.8 | 20.6% | 73.4% | estimated ± 1.0 pp, medium confidence |
| Claude Sonnet 4.6 | 8.3% | 62.8% | estimated ± 1.0 pp, low confidence |
| Gemini 3.7 Flash | 47.9% | 78.5% | estimated ± 1.0 pp, medium confidence |
| Gemini 3.8 Flash | 59.0% | 79.3% | estimated ± 1.0 pp, medium confidence |
| Gemini 4 Argon | 69.2% | 79.8% | estimated ± 1.0 pp, medium confidence |
| GPT-5.5 | 13.0% | 68.8% | estimated ± 1.0 pp, low confidence |
| GPT-5.6 Luna | 45.6% | 78.3% | estimated ± 1.0 pp, medium confidence |
| GPT-5.6 Sol | 62.6% | 79.5% | estimated ± 1.0 pp, medium confidence |
| GPT-5.6 Terra | 50.2% | 78.7% | estimated ± 1.0 pp, medium confidence |
| GPT-6 Astra | 72.6% | 80.0% | estimated ± 1.0 pp, low confidence |
| GPT-6 Sol | 60.5% | 79.4% | estimated ± 1.0 pp, medium confidence |
| Kimi K2.6 | 4.6% | 52.5% | estimated ± 1.0 pp, low confidence |
| MiniMax M3 | 4.6% | 52.5% | estimated ± 1.0 pp, low confidence |
| Muse Spark 1.1 | 14.2% | 69.8% | estimated ± 1.0 pp, low confidence |
| Muse Spark 1.3 | 66.9% | 79.7% | estimated ± 1.0 pp, medium confidence |
| Qwen3.7 Plus | 2.8% | 42.5% | estimated ± 1.0 pp, low confidence |