benchgap
Calibration

Gert Labs → ApprenticeBench

ApprenticeBench is estimated from Gert Labs with a linear curve fitted on 5 models measured on both: y = 1.4971·x + -0.8927, R² = 0.87, cross-validated error 3.1 pp. It is used for 47 estimates.

Estimated modelGert LabsApprenticeBenchSource
Claude 4 Sonnet39.7%0.0%estimated ± 3.1 pp, low confidence
Claude Opus 4.564.2%6.9%estimated ± 3.1 pp, medium confidence
Claude Opus 4.873.0%20.0%estimated ± 3.1 pp, low confidence
Claude Sonnet 4.548.5%0.0%estimated ± 3.1 pp, low confidence
DeepSeek V3.229.6%0.0%estimated ± 3.1 pp, low confidence
Gemini 2.5 Pro42.0%0.0%estimated ± 3.1 pp, low confidence
Gemini 3.1 Flash-Lite38.5%0.0%estimated ± 3.1 pp, low confidence
Gemini 3.1 Pro56.9%0.0%estimated ± 3.1 pp, low confidence
Gemini 3.5 Flash61.9%3.3%estimated ± 3.1 pp, medium confidence
Gemini 3 Flash56.6%0.0%estimated ± 3.1 pp, low confidence
Gemini 3 Pro63.2%5.4%estimated ± 3.1 pp, medium confidence
Gemma 4 31B35.3%0.0%estimated ± 3.1 pp, low confidence
GLM-4.740.0%0.0%estimated ± 3.1 pp, low confidence
GLM-551.0%0.0%estimated ± 3.1 pp, low confidence
GLM-5.160.1%0.7%estimated ± 3.1 pp, low confidence
GLM-5V-Turbo30.8%0.0%estimated ± 3.1 pp, low confidence
GPT-4.125.7%0.0%estimated ± 3.1 pp, low confidence
GPT-5.141.2%0.0%estimated ± 3.1 pp, low confidence
GPT-5.1-Codex49.7%0.0%estimated ± 3.1 pp, low confidence
GPT-5.246.5%0.0%estimated ± 3.1 pp, low confidence
GPT-5.2-Codex51.8%0.0%estimated ± 3.1 pp, low confidence
GPT-5.3 Codex57.5%0.0%estimated ± 3.1 pp, low confidence
GPT-OSS 120B29.6%0.0%estimated ± 3.1 pp, low confidence
Grok 442.3%0.0%estimated ± 3.1 pp, low confidence
Grok 4.1 Fast47.3%0.0%estimated ± 3.1 pp, low confidence
Grok 4.2038.4%0.0%estimated ± 3.1 pp, low confidence
Grok 4.343.9%0.0%estimated ± 3.1 pp, low confidence
Grok Build 0.149.2%0.0%estimated ± 3.1 pp, low confidence
Hy3 Preview36.9%0.0%estimated ± 3.1 pp, low confidence
Kimi K2.656.8%0.0%estimated ± 3.1 pp, low confidence
Kimi K2.545.9%0.0%estimated ± 3.1 pp, low confidence
Kimi K2.5 (Reasoning)32.6%0.0%estimated ± 3.1 pp, low confidence
MiMo-V2.546.9%0.0%estimated ± 3.1 pp, low confidence
MiMo-V2.5-Pro62.7%4.6%estimated ± 3.1 pp, medium confidence
MiMo-V2-Pro36.7%0.0%estimated ± 3.1 pp, low confidence
MiniMax M2.740.4%0.0%estimated ± 3.1 pp, low confidence
Mistral Medium 3.5 128B39.1%0.0%estimated ± 3.1 pp, low confidence
Qwen3.5-27B39.4%0.0%estimated ± 3.1 pp, low confidence
Qwen3.5-35B-A3B29.0%0.0%estimated ± 3.1 pp, low confidence
Qwen3.5 397B46.8%0.0%estimated ± 3.1 pp, low confidence
Qwen3.6-27B54.8%0.0%estimated ± 3.1 pp, low confidence
Qwen3.6-35B-A3B42.7%0.0%estimated ± 3.1 pp, low confidence
Qwen3.6 Plus50.6%0.0%estimated ± 3.1 pp, low confidence
Qwen3.7 Max64.3%7.0%estimated ± 3.1 pp, medium confidence
Qwen3 Max43.7%0.0%estimated ± 3.1 pp, low confidence
Step 3.7 Flash51.6%0.0%estimated ± 3.1 pp, low confidence
Trinity-Large-Thinking32.6%0.0%estimated ± 3.1 pp, low confidence