benchgap
Calibration

ApprenticeBench → Toolathlon

Toolathlon is estimated from ApprenticeBench with a linear curve fitted on 5 models measured on both: y = 0.2198·x + 0.5142, R² = 0.68, cross-validated error 1.7 pp. It is used for 13 estimates.

Estimated modelApprenticeBenchToolathlonSource
Claude Fable 534.0%58.9%estimated ± 1.7 pp, low confidence
Claude Fable 5.172.0%67.2%estimated ± 1.7 pp, low confidence
Claude Opus 4.65.0%52.5%estimated ± 1.7 pp, low confidence
Claude Opus 4.77.0%53.0%estimated ± 1.7 pp, medium confidence
Claude Opus 536.0%59.3%estimated ± 1.7 pp, low confidence
Claude Sonnet 4.62.0%51.9%estimated ± 1.7 pp, low confidence
Claude Sonnet 516.0%54.9%estimated ± 1.7 pp, medium confidence
Gemini 3.7 Flash16.0%54.9%estimated ± 1.7 pp, medium confidence
Gemini 3.8 Flash24.0%56.7%estimated ± 1.7 pp, medium confidence
GPT-6 Astra68.0%66.4%estimated ± 1.7 pp, low confidence
Grok 4.613.0%54.3%estimated ± 1.7 pp, medium confidence
Kimi K318.0%55.4%estimated ± 1.7 pp, medium confidence
Muse Spark 1.319.0%55.6%estimated ± 1.7 pp, medium confidence