benchgap
Calibration

ApprenticeBench → Gert Labs

Gert Labs is estimated from ApprenticeBench with a linear curve fitted on 5 models measured on both: y = 0.5824·x + 0.6039, R² = 0.87, cross-validated error 3.0 pp. It is used for 13 estimates.

Estimated modelApprenticeBenchGert LabsSource
Claude Fable 534.0%80.2%estimated ± 3.0 pp, low confidence
Claude Fable 5.172.0%100.0%estimated ± 3.0 pp, low confidence
Claude Opus 536.0%81.4%estimated ± 3.0 pp, low confidence
Claude Sonnet 516.0%69.7%estimated ± 3.0 pp, medium confidence
Gemini 3.7 Flash16.0%69.7%estimated ± 3.0 pp, medium confidence
Gemini 3.8 Flash24.0%74.4%estimated ± 3.0 pp, low confidence
GPT-5.6 Luna7.0%64.5%estimated ± 3.0 pp, medium confidence
GPT-5.6 Sol26.0%75.5%estimated ± 3.0 pp, low confidence
GPT-5.6 Terra16.0%69.7%estimated ± 3.0 pp, medium confidence
GPT-6 Astra68.0%100.0%estimated ± 3.0 pp, low confidence
Grok 4.613.0%68.0%estimated ± 3.0 pp, medium confidence
Kimi K318.0%70.9%estimated ± 3.0 pp, medium confidence
Muse Spark 1.319.0%71.5%estimated ± 3.0 pp, medium confidence