benchgap
Calibration

ApprenticeBench → OSWorld-Verified

OSWorld-Verified is estimated from ApprenticeBench with a linear curve fitted on 6 models measured on both: y = 0.4175·x + 0.7133, R² = 0.90, cross-validated error 1.9 pp. It is used for 12 estimates.

Estimated modelApprenticeBenchOSWorld-VerifiedSource
Claude Fable 5.172.0%100.0%estimated ± 1.9 pp, low confidence
Claude Opus 4.77.0%74.2%estimated ± 1.9 pp, medium confidence
Claude Opus 536.0%86.4%estimated ± 1.9 pp, low confidence
Gemini 3.7 Flash16.0%78.0%estimated ± 1.9 pp, medium confidence
Gemini 3.8 Flash24.0%81.3%estimated ± 1.9 pp, medium confidence
GPT-5.6 Luna7.0%74.2%estimated ± 1.9 pp, medium confidence
GPT-5.6 Sol26.0%82.2%estimated ± 1.9 pp, medium confidence
GPT-5.6 Terra16.0%78.0%estimated ± 1.9 pp, medium confidence
GPT-6 Astra68.0%99.7%estimated ± 1.9 pp, low confidence
Grok 4.613.0%76.8%estimated ± 1.9 pp, medium confidence
Kimi K318.0%78.8%estimated ± 1.9 pp, medium confidence
Muse Spark 1.319.0%79.3%estimated ± 1.9 pp, medium confidence