benchgap
Agentic · tools

JobBench leaderboard

As of 2026-10-07, the highest measured score on JobBench is 64.9% by Muse Spark 1.3. 150 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 569.8%estimated ± 3.4 pp, low confidence
2GLM-5.3-Flash66.2%estimated ± 3.4 pp, low confidence
3GPT-6 Astra66.0%estimated ± 4.9 pp, low confidence
4Gemini 4 Argon65.7%estimated ± 4.9 pp, low confidence
5Grok 4.765.4%estimated ± 8.6 pp, low confidence
6Claude Fable 5.165.2%estimated ± 3.4 pp, low confidence
7Claude Opus 5.565.2%estimated ± 3.4 pp, low confidence
8Claude Sonnet 5.565.2%estimated ± 3.4 pp, low confidence
9GPT-5.6 Sol65.0%estimated ± 4.9 pp, medium confidence
10Muse Spark 1.364.9%measured
11GPT-6 Sol64.8%estimated ± 4.9 pp, medium confidence
12Gemini 3.8 Flash64.6%estimated ± 4.9 pp, medium confidence
13GPT-5.6 Terra63.3%estimated ± 4.9 pp, medium confidence
14Gemini 3.7 Flash62.9%estimated ± 4.9 pp, medium confidence
15GPT-5.6 Luna62.5%estimated ± 4.9 pp, medium confidence
16MiMo-V2.6-Pro62.0%measured
17Hy4 preview61.7%measured
18MiMo-V2.6-Flash61.2%measured
19Grok 4.659.5%estimated ± 8.6 pp, medium confidence
20Ling 3.1 Flash59.5%estimated ± 8.6 pp, medium confidence
21Step 5 Preview59.0%measured
22Claude Fable 558.9%estimated ± 8.6 pp, medium confidence
23DeepSeek V4 Pro 081358.3%estimated ± 3.4 pp, low confidence
24DeepSeek V4.1 Flash58.2%estimated ± 8.6 pp, medium confidence
25GPT-6.1 Sol56.7%estimated ± 8.6 pp, medium confidence
26GLM-5.356.1%estimated ± 3.4 pp, low confidence
27Qwen3.8 Max Preview55.9%estimated ± 8.2 pp, low confidence
28Qwen3.8-Flash-Next55.7%measured
29Muse Spark 1.154.7%measured
30Gemini 3.5 Flash54.3%estimated ± 6.8 pp, medium confidence
31Fugu Cyber53.9%estimated ± 8.6 pp, low confidence
32Gemini 3.8 Flash Cyber53.5%estimated ± 8.6 pp, low confidence
33Claude Opus 4.853.4%estimated ± 4.9 pp, medium confidence
34Qwen3.8 Max53.4%measured
35Kimi K352.9%measured
36Claude Mythos 552.4%estimated ± 8.6 pp, low confidence
37Ornith-1.5-397B52.3%estimated ± 3.4 pp, low confidence
38Gemini 3.5 Flash Cyber52.2%estimated ± 8.6 pp, low confidence
39Claude Mythos Preview52.1%estimated ± 8.6 pp, low confidence
40GPT-5.5 Pro51.8%estimated ± 10.1 pp, low confidence
41Holo3-35B-A3B51.8%estimated ± 9.0 pp, medium confidence
42GPT-5.4 Pro50.8%estimated ± 10.1 pp, low confidence
43Muse Spark 1.250.6%estimated ± 8.6 pp, medium confidence
44DeepSeek V4 Flash 073150.4%estimated ± 3.4 pp, low confidence
45Atria Dawn Preview50.3%measured
46Claude Sonnet 549.9%estimated ± 8.6 pp, medium confidence
47Beam49.1%estimated ± 6.8 pp, medium confidence
48Holo3-122B-A10B49.0%estimated ± 9.0 pp, medium confidence
49GPT-6 Luna48.1%estimated ± 8.6 pp, medium confidence
50Claude Opus 4.747.3%estimated ± 4.9 pp, medium confidence
51Mistral Large 447.3%estimated ± 8.6 pp, medium confidence
52GLM-5.247.1%estimated ± 6.8 pp, medium confidence
53Qwen3.7 Max46.7%estimated ± 6.8 pp, medium confidence
54Kimi K2.7 Code46.4%estimated ± 6.8 pp, medium confidence
55Claude Opus 4.7 (Adaptive)45.9%measured
56Muse Glimmer 30B45.9%estimated ± 6.8 pp, medium confidence
57K-EXAONE 2.045.9%estimated ± 8.4 pp, low confidence
58Ornith-1.0-397B45.5%estimated ± 8.4 pp, low confidence
59Grok 4.545.2%estimated ± 8.6 pp, medium confidence
60UI-Mate-27B45.0%estimated ± 9.0 pp, medium confidence
61Inkling44.5%estimated ± 6.8 pp, medium confidence
62GPT-5.542.7%measured
63GLM-5.142.4%estimated ± 6.8 pp, medium confidence
64Qwen3.6-27B42.3%estimated ± 8.4 pp, low confidence
65Claude Opus 4.6 (Adaptive)41.4%estimated ± 8.2 pp, low confidence
66Ornith-1.0-35B40.1%estimated ± 8.4 pp, low confidence
67Gemini 3.1 Pro39.9%estimated ± 8.2 pp, low confidence
68GPT-5.438.9%measured
69Gemini 3.6 Flash38.7%estimated ± 8.6 pp, medium confidence
70Apodex 1.138.7%estimated ± 8.2 pp, low confidence
71Apodex 1.1 Mini38.7%estimated ± 8.2 pp, low confidence
72Claude Sonnet 4.636.9%measured
73Claude Opus 4.636.7%measured
74Step 3.7 Flash36.0%estimated ± 7.3 pp, low confidence
75Hy3 Preview34.4%estimated ± 8.6 pp, medium confidence
76GPT-5.234.3%measured
77MiniMax M2.734.1%estimated ± 7.3 pp, low confidence
78Agents-A133.9%estimated ± 10.1 pp, low confidence
79GPT-5.3 Codex33.7%measured
80Muse Spark33.6%estimated ± 8.4 pp, low confidence
81Solar Pro 433.5%estimated ± 6.8 pp, medium confidence
82Qwen3.8-27B33.4%measured
83UI-Mate-9B32.8%estimated ± 9.0 pp, medium confidence
84Ornith-1.0-9B32.8%estimated ± 8.4 pp, low confidence
85Qwen3.5-27B32.8%estimated ± 9.0 pp, medium confidence
86Qwen3.5-35B-A3B32.8%estimated ± 9.0 pp, medium confidence
87LFM2.5-2.6B32.5%estimated ± 8.4 pp, low confidence
88Quasar 438B32.4%estimated ± 8.6 pp, medium confidence
89Claude Opus 4.532.3%measured
90MiMo-V2.531.8%estimated ± 8.4 pp, low confidence
91Ling 3.0 Flash VL31.1%estimated ± 8.6 pp, medium confidence
92Solar Open 231.0%estimated ± 6.8 pp, medium confidence
93Nemotron 3 Ultra31.0%estimated ± 8.6 pp, medium confidence
94GPT-5.4 mini30.7%estimated ± 6.8 pp, medium confidence
95GPT-5.4 nano29.5%estimated ± 6.8 pp, medium confidence
96Kimi K2.627.7%estimated ± 4.9 pp, low confidence
97MiniMax M327.7%estimated ± 4.9 pp, low confidence
98Claude Sonnet 4.527.7%measured
99GPT-5.1-Codex26.2%measured
100GPT-5.2-Codex26.0%measured
101MiMo-V2-Pro25.7%estimated ± 8.4 pp, low confidence
102Hy325.1%estimated ± 8.6 pp, medium confidence
103LLaDA2.2-mini24.8%estimated ± 8.4 pp, low confidence
104Agents-A1-4B24.7%estimated ± 10.1 pp, low confidence
105GLM-5-Turbo22.8%estimated ± 8.4 pp, low confidence
106LLaDA2.2-flash22.7%estimated ± 6.8 pp, medium confidence
107LongCat-Flash-Lite-Sparse22.3%estimated ± 6.8 pp, medium confidence
108GLM-4.722.0%estimated ± 8.6 pp, medium confidence
109Claude 4.1 Opus21.9%measured
110Gemini 3.5 Flash-Lite20.1%estimated ± 8.6 pp, medium confidence
111GLM-5V-Turbo20.0%estimated ± 8.4 pp, low confidence
112dots3-note Preview19.9%estimated ± 3.4 pp, low confidence
113Qwen3.7 Plus19.9%estimated ± 4.9 pp, low confidence
114A.X K218.8%estimated ± 8.6 pp, medium confidence
115Qwen3.6 Plus18.5%estimated ± 5.1 pp, low confidence
116Qwen3.5 Plus18.5%measured
117Claude 4 Sonnet18.4%measured
118Inkling-Small17.9%estimated ± 3.4 pp, low confidence
119Ling 3.0 Flash FP817.8%estimated ± 8.6 pp, medium confidence
120Qwen3.5 397B17.2%estimated ± 5.1 pp, low confidence
121Grok 4.316.8%estimated ± 8.2 pp, low confidence
122Claude Haiku 4.516.0%measured
123Ling 3.0 Flash15.4%estimated ± 5.1 pp, low confidence
124Gemini 3 Flash11.4%measured
125Gemini 3 Pro11.4%measured
126Laguna S 2.111.3%estimated ± 3.4 pp, low confidence
127GPT-5.110.4%estimated ± 8.6 pp, low confidence
128Ornith-1.5-35B-A3B10.1%estimated ± 3.4 pp, low confidence
129Qwen3.5-122B-A10B9.6%estimated ± 8.6 pp, low confidence
130MiMo-V2-Omni9.2%estimated ± 8.4 pp, low confidence
131Kimi K2.58.7%measured
132GPT-5 (high)8.5%measured
133Kimi K2.5 (Reasoning)8.3%estimated ± 8.2 pp, low confidence
134Mistral Medium 3.5 128B6.4%estimated ± 8.6 pp, low confidence
135DeepSeek V3.25.0%estimated ± 8.4 pp, low confidence
136Ornith-1.5-9B4.0%estimated ± 3.4 pp, low confidence
137MiniCPM5-2B3.0%estimated ± 8.6 pp, low confidence
138Claude Haiku 5.50.7%estimated ± 9.1 pp, low confidence
139Celeris-10.0%estimated ± 8.6 pp, low confidence
140Command A+0.0%estimated ± 8.6 pp, low confidence
141DeepSeek V30.0%estimated ± 8.6 pp, low confidence
142DeepSeek V3 03240.0%estimated ± 8.6 pp, low confidence
143Gemini 2.5 Pro0.0%estimated ± 8.6 pp, low confidence
144Gemma 3 27B0.0%estimated ± 8.6 pp, low confidence
145Gemma 4 12B0.0%estimated ± 8.6 pp, low confidence
146Gemma 4 26B A4B0.0%estimated ± 8.6 pp, low confidence
147Gemma 4 31B0.0%estimated ± 8.6 pp, low confidence
148Gemma 4 E2B0.0%estimated ± 8.6 pp, low confidence
149Gemma 4 E4B0.0%estimated ± 8.6 pp, low confidence
150GLM-50.0%estimated ± 5.1 pp, low confidence
151GPT-4.1 mini0.0%estimated ± 8.6 pp, low confidence
152GPT-4.1 nano0.0%estimated ± 8.6 pp, low confidence
153GPT-4o0.0%estimated ± 8.6 pp, low confidence
154GPT-4o mini0.0%estimated ± 8.6 pp, low confidence
155GPT-OSS 120B0.0%estimated ± 8.2 pp, low confidence
156GPT-OSS 20B0.0%estimated ± 8.2 pp, low confidence
157Granite 4.2 30B0.0%estimated ± 8.6 pp, low confidence
158Granite 4.2 3B0.0%estimated ± 8.6 pp, low confidence
159Granite 4.2 8B0.0%estimated ± 8.6 pp, low confidence
160K-Exaone0.0%estimated ± 8.6 pp, low confidence
161Ling 2.6 Flash0.0%estimated ± 8.6 pp, low confidence
162Ling 3.0 Tiny0.0%estimated ± 8.6 pp, low confidence
163Llama 4 Maverick0.0%estimated ± 8.6 pp, low confidence
164Llama 4 Scout0.0%estimated ± 8.6 pp, low confidence
165Mercury 2.50.0%estimated ± 8.6 pp, low confidence
166MiMo-V2.5-Pro0.0%estimated ± 8.2 pp, low confidence
167MiMo-V2-Flash0.0%estimated ± 8.6 pp, low confidence
168Mistral Large 30.0%estimated ± 8.6 pp, low confidence
169Mistral Small 40.0%estimated ± 8.6 pp, low confidence
170Mistral Small 4 (Reasoning)0.0%estimated ± 8.6 pp, low confidence
171Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 8.6 pp, low confidence
172Nemotron 3 Nano 30B0.0%estimated ± 8.6 pp, low confidence
173Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 8.6 pp, low confidence
174Nemotron 3 Super 100B0.0%estimated ± 8.2 pp, low confidence
175North Mini Code0.0%estimated ± 8.6 pp, low confidence
176Qwen3.6-35B-A3B0.0%estimated ± 5.1 pp, low confidence
177Solar Pro 30.0%estimated ± 8.6 pp, low confidence
178Trinity-Large-Preview0.0%estimated ± 8.6 pp, low confidence
179Trinity-Large-Thinking0.0%estimated ± 8.6 pp, low confidence
180Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 8.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General