Calibration
ApprenticeBench → OSWorld-Verified
OSWorld-Verified is estimated from ApprenticeBench with a linear curve fitted on 6 models measured on both: y = 0.4175·x + 0.7133, R² = 0.90, cross-validated error 1.9 pp. It is used for 12 estimates.
| Estimated model | ApprenticeBench | OSWorld-Verified | Source |
|---|---|---|---|
| Claude Fable 5.1 | 72.0% | 100.0% | estimated ± 1.9 pp, low confidence |
| Claude Opus 4.7 | 7.0% | 74.2% | estimated ± 1.9 pp, medium confidence |
| Claude Opus 5 | 36.0% | 86.4% | estimated ± 1.9 pp, low confidence |
| Gemini 3.7 Flash | 16.0% | 78.0% | estimated ± 1.9 pp, medium confidence |
| Gemini 3.8 Flash | 24.0% | 81.3% | estimated ± 1.9 pp, medium confidence |
| GPT-5.6 Luna | 7.0% | 74.2% | estimated ± 1.9 pp, medium confidence |
| GPT-5.6 Sol | 26.0% | 82.2% | estimated ± 1.9 pp, medium confidence |
| GPT-5.6 Terra | 16.0% | 78.0% | estimated ± 1.9 pp, medium confidence |
| GPT-6 Astra | 68.0% | 99.7% | estimated ± 1.9 pp, low confidence |
| Grok 4.6 | 13.0% | 76.8% | estimated ± 1.9 pp, medium confidence |
| Kimi K3 | 18.0% | 78.8% | estimated ± 1.9 pp, medium confidence |
| Muse Spark 1.3 | 19.0% | 79.3% | estimated ± 1.9 pp, medium confidence |