benchgap
Agentic · tools

ApprenticeBench leaderboard

As of 2026-10-07, the highest measured score on ApprenticeBench is 72.0% by Claude Fable 5.1. 94 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.172.0%measured
2GPT-6 Astra68.0%measured
3Claude Opus 536.0%measured
4Claude Fable 534.0%measured
5Qwen3.8 Max33.3%estimated ± 5.2 pp, low confidence
6MiMo-V2.6-Flash32.7%estimated ± 4.0 pp, low confidence
7MiMo-V2.6-Pro32.6%estimated ± 4.0 pp, low confidence
8DeepSeek V4.1 Flash31.0%estimated ± 4.0 pp, low confidence
9Ling 3.1 Flash30.9%estimated ± 4.0 pp, low confidence
10Fugu Cyber30.0%estimated ± 4.0 pp, low confidence
11Atria Dawn Preview29.5%estimated ± 4.0 pp, low confidence
12Qwen3.8-27B29.5%estimated ± 5.2 pp, low confidence
13Gemini 3.8 Flash Cyber29.1%estimated ± 4.0 pp, low confidence
14Gemini 3.6 Flash26.7%estimated ± 5.2 pp, low confidence
15Step 5 Preview26.4%estimated ± 4.0 pp, low confidence
16GPT-5.6 Sol26.0%measured
17GLM-5.325.9%estimated ± 4.0 pp, medium confidence
18Holo3-35B-A3B25.7%estimated ± 5.2 pp, low confidence
19Claude Mythos 524.2%estimated ± 4.0 pp, medium confidence
20Gemini 3.8 Flash24.0%measured
21dots3-note Preview23.1%estimated ± 10.1 pp, low confidence
22DeepSeek V4 Pro 081322.8%estimated ± 4.0 pp, medium confidence
23Gemini 3.5 Flash Cyber22.5%estimated ± 4.0 pp, medium confidence
24Claude Mythos Preview22.2%estimated ± 4.0 pp, medium confidence
25Qwen3.8 Max Preview20.3%estimated ± 5.0 pp, low confidence
26GPT-5.520.0%measured
27Claude Opus 4.820.0%estimated ± 3.1 pp, low confidence
28Qwen3.8-Flash-Next19.0%estimated ± 5.6 pp, low confidence
29Muse Spark 1.319.0%measured
30Kimi K318.0%measured
31Holo3-122B-A10B17.7%estimated ± 5.2 pp, low confidence
32Claude Sonnet 516.0%measured
33Gemini 3.7 Flash16.0%measured
34GPT-5.6 Terra16.0%measured
35UI-Mate-27B13.7%estimated ± 5.2 pp, low confidence
36Grok 4.613.0%measured
37Beam11.8%estimated ± 10.1 pp, low confidence
38GPT-5.411.0%measured
39GLM-5.210.0%estimated ± 5.0 pp, low confidence
40Claude Opus 4.6 (Adaptive)9.2%estimated ± 5.0 pp, low confidence
41Hy4 preview8.8%estimated ± 4.0 pp, medium confidence
42Gemini 3.5 Flash-Lite7.2%estimated ± 5.2 pp, low confidence
43Apodex 1.17.0%estimated ± 5.0 pp, low confidence
44Apodex 1.1 Mini7.0%estimated ± 5.0 pp, low confidence
45Claude Opus 4.77.0%measured
46GPT-5.6 Luna7.0%measured
47Qwen3.7 Max7.0%estimated ± 3.1 pp, medium confidence
48Claude Opus 4.56.9%estimated ± 3.1 pp, medium confidence
49DeepSeek V4 Flash 07316.2%estimated ± 4.0 pp, medium confidence
50Gemini 3 Pro5.4%estimated ± 3.1 pp, medium confidence
51Claude Opus 4.65.0%measured
52MiMo-V2.5-Pro4.6%estimated ± 3.1 pp, medium confidence
53Claude Opus 4.7 (Adaptive)4.0%estimated ± 4.0 pp, medium confidence
54GPT-5.4 mini3.5%estimated ± 5.0 pp, low confidence
55Muse Spark 1.13.4%estimated ± 4.0 pp, low confidence
56Muse Spark3.4%estimated ± 4.0 pp, low confidence
57Gemini 3.5 Flash3.3%estimated ± 3.1 pp, medium confidence
58Claude Sonnet 4.62.0%measured
59Mercury 2.51.3%estimated ± 10.1 pp, low confidence
60GLM-5.10.7%estimated ± 3.1 pp, low confidence
61Claude 4.1 Opus0.5%estimated ± 5.6 pp, low confidence
62Qwen3.5 Plus0.2%estimated ± 5.6 pp, low confidence
63Claude Haiku 4.50.1%estimated ± 5.6 pp, low confidence
64GPT-5 (high)0.0%estimated ± 5.6 pp, low confidence
65Claude 4 Sonnet0.0%estimated ± 3.1 pp, low confidence
66Claude Sonnet 4.50.0%estimated ± 3.1 pp, low confidence
67DeepSeek V3.20.0%estimated ± 3.1 pp, low confidence
68Gemini 2.5 Pro0.0%estimated ± 3.1 pp, low confidence
69Gemini 3.1 Flash-Lite0.0%estimated ± 3.1 pp, low confidence
70Gemini 3.1 Pro0.0%estimated ± 3.1 pp, low confidence
71Gemini 3 Flash0.0%estimated ± 3.1 pp, low confidence
72Gemma 4 31B0.0%estimated ± 3.1 pp, low confidence
73GLM-4.70.0%estimated ± 3.1 pp, low confidence
74GLM-50.0%estimated ± 3.1 pp, low confidence
75GLM-5V-Turbo0.0%estimated ± 3.1 pp, low confidence
76GPT-4.10.0%estimated ± 3.1 pp, low confidence
77GPT-5.10.0%estimated ± 3.1 pp, low confidence
78GPT-5.1-Codex0.0%estimated ± 3.1 pp, low confidence
79GPT-5.20.0%estimated ± 3.1 pp, low confidence
80GPT-5.2-Codex0.0%estimated ± 3.1 pp, low confidence
81GPT-5.3 Codex0.0%estimated ± 3.1 pp, low confidence
82GPT-5.4 nano0.0%estimated ± 5.0 pp, low confidence
83GPT-OSS 120B0.0%estimated ± 3.1 pp, low confidence
84GPT-OSS 20B0.0%estimated ± 5.0 pp, low confidence
85Grok 40.0%estimated ± 3.1 pp, low confidence
86Grok 4.1 Fast0.0%estimated ± 3.1 pp, low confidence
87Grok 4.200.0%estimated ± 3.1 pp, low confidence
88Grok 4.30.0%estimated ± 3.1 pp, low confidence
89Grok Build 0.10.0%estimated ± 3.1 pp, low confidence
90Hy3 Preview0.0%estimated ± 3.1 pp, low confidence
91Kimi K2.60.0%estimated ± 3.1 pp, low confidence
92Kimi K2.50.0%estimated ± 3.1 pp, low confidence
93Kimi K2.5 (Reasoning)0.0%estimated ± 3.1 pp, low confidence
94MiMo-V2.50.0%estimated ± 3.1 pp, low confidence
95MiMo-V2-Pro0.0%estimated ± 3.1 pp, low confidence
96MiniMax M2.70.0%estimated ± 3.1 pp, low confidence
97MiniMax M30.0%estimated ± 5.2 pp, low confidence
98Mistral Medium 3.5 128B0.0%estimated ± 3.1 pp, low confidence
99Muse Glimmer 30B0.0%estimated ± 5.2 pp, low confidence
100Nemotron 3 Super 100B0.0%estimated ± 5.0 pp, low confidence
101Qwen3.5-122B-A10B0.0%estimated ± 5.2 pp, low confidence
102Qwen3.5-27B0.0%estimated ± 3.1 pp, low confidence
103Qwen3.5-35B-A3B0.0%estimated ± 3.1 pp, low confidence
104Qwen3.5 397B0.0%estimated ± 3.1 pp, low confidence
105Qwen3.6-27B0.0%estimated ± 3.1 pp, low confidence
106Qwen3.6-35B-A3B0.0%estimated ± 3.1 pp, low confidence
107Qwen3.6 Plus0.0%estimated ± 3.1 pp, low confidence
108Qwen3.7 Plus0.0%estimated ± 5.0 pp, low confidence
109Qwen3 Max0.0%estimated ± 3.1 pp, low confidence
110Step 3.7 Flash0.0%estimated ± 3.1 pp, low confidence
111Trinity-Large-Thinking0.0%estimated ± 3.1 pp, low confidence
112UI-Mate-9B0.0%estimated ± 5.2 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General