Calibration
ApprenticeBench → Gert Labs
Gert Labs is estimated from ApprenticeBench with a linear curve fitted on 5 models measured on both: y = 0.5824·x + 0.6039, R² = 0.87, cross-validated error 3.0 pp. It is used for 13 estimates.
| Estimated model | ApprenticeBench | Gert Labs | Source |
|---|---|---|---|
| Claude Fable 5 | 34.0% | 80.2% | estimated ± 3.0 pp, low confidence |
| Claude Fable 5.1 | 72.0% | 100.0% | estimated ± 3.0 pp, low confidence |
| Claude Opus 5 | 36.0% | 81.4% | estimated ± 3.0 pp, low confidence |
| Claude Sonnet 5 | 16.0% | 69.7% | estimated ± 3.0 pp, medium confidence |
| Gemini 3.7 Flash | 16.0% | 69.7% | estimated ± 3.0 pp, medium confidence |
| Gemini 3.8 Flash | 24.0% | 74.4% | estimated ± 3.0 pp, low confidence |
| GPT-5.6 Luna | 7.0% | 64.5% | estimated ± 3.0 pp, medium confidence |
| GPT-5.6 Sol | 26.0% | 75.5% | estimated ± 3.0 pp, low confidence |
| GPT-5.6 Terra | 16.0% | 69.7% | estimated ± 3.0 pp, medium confidence |
| GPT-6 Astra | 68.0% | 100.0% | estimated ± 3.0 pp, low confidence |
| Grok 4.6 | 13.0% | 68.0% | estimated ± 3.0 pp, medium confidence |
| Kimi K3 | 18.0% | 70.9% | estimated ± 3.0 pp, medium confidence |
| Muse Spark 1.3 | 19.0% | 71.5% | estimated ± 3.0 pp, medium confidence |