benchgap
Knowledge & reasoning

MMLU-Pro (Arcee) leaderboard

As of 2026-10-07, the highest measured score on MMLU-Pro (Arcee) is 89.1% by Claude Opus 4.6. 220 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5100.0%estimated ± 4.0 pp, low confidence
2Claude Fable 5.1100.0%estimated ± 4.0 pp, low confidence
3Claude Haiku 5.5100.0%estimated ± 4.0 pp, low confidence
4Claude Opus 4.6 (Adaptive)100.0%estimated ± 4.0 pp, low confidence
5Claude Opus 5100.0%estimated ± 4.0 pp, low confidence
6Claude Opus 5.5100.0%estimated ± 4.0 pp, low confidence
7Claude Sonnet 5100.0%estimated ± 4.0 pp, low confidence
8Claude Sonnet 5.5100.0%estimated ± 4.0 pp, low confidence
9Gemini 3.6 Flash100.0%estimated ± 4.0 pp, low confidence
10Gemini 3.7 Flash100.0%estimated ± 4.0 pp, low confidence
11Gemini 3.8 Flash100.0%estimated ± 4.0 pp, low confidence
12Gemini 3 Pro100.0%estimated ± 4.0 pp, low confidence
13Gemini 3 Pro Deep Think100.0%estimated ± 4.0 pp, low confidence
14Gemini 4 Argon100.0%estimated ± 4.0 pp, low confidence
15GLM-5.3100.0%estimated ± 4.0 pp, low confidence
16GLM-5.3-Flash100.0%estimated ± 4.0 pp, low confidence
17GPT-5.2100.0%estimated ± 4.0 pp, low confidence
18GPT-5.2-Codex100.0%estimated ± 4.0 pp, low confidence
19GPT-5.3 Codex100.0%estimated ± 4.0 pp, low confidence
20GPT-5.4 mini100.0%estimated ± 4.0 pp, low confidence
21GPT-5.4 nano100.0%estimated ± 4.0 pp, low confidence
22GPT-5.4 Pro100.0%estimated ± 4.0 pp, low confidence
23GPT-5.5 Pro100.0%estimated ± 4.0 pp, low confidence
24GPT-6.1 Sol100.0%estimated ± 4.0 pp, low confidence
25GPT-6 Luna100.0%estimated ± 4.0 pp, low confidence
26GPT-6 Sol100.0%estimated ± 4.0 pp, low confidence
27Grok 4.3100.0%estimated ± 4.0 pp, low confidence
28Grok 4.5100.0%estimated ± 4.0 pp, low confidence
29Grok 4.6100.0%estimated ± 4.0 pp, low confidence
30Grok 4.7100.0%estimated ± 4.0 pp, low confidence
31Kimi K2.7 Code100.0%estimated ± 4.0 pp, low confidence
32Ling 3.1 Flash100.0%estimated ± 4.0 pp, low confidence
33MiMo-V2.6-Flash100.0%estimated ± 4.0 pp, low confidence
34MiMo-V2.6-Pro100.0%estimated ± 4.0 pp, low confidence
35Mistral Large 4100.0%estimated ± 4.0 pp, low confidence
36Muse Spark 1.1100.0%estimated ± 4.0 pp, low confidence
37Muse Spark 1.2100.0%estimated ± 4.0 pp, low confidence
38Muse Spark 1.3100.0%estimated ± 4.0 pp, low confidence
39Quasar 438B100.0%estimated ± 4.0 pp, low confidence
40Qwen3.8 Max Preview100.0%estimated ± 4.0 pp, low confidence
41GPT-5.1-Codex98.5%estimated ± 4.0 pp, low confidence
42GPT-5.1-Codex-Max98.5%estimated ± 4.0 pp, low confidence
43GPT-5 (high)98.5%estimated ± 4.0 pp, low confidence
44Claude Opus 4.796.3%estimated ± 4.0 pp, low confidence
45GPT-5.195.5%estimated ± 4.0 pp, low confidence
46Hy395.5%estimated ± 4.0 pp, low confidence
47Apodex 1.194.4%estimated ± 4.0 pp, low confidence
48Apodex 1.1 Mini94.4%estimated ± 4.0 pp, low confidence
49Claude Opus 4.5 Thinking94.4%estimated ± 4.0 pp, low confidence
50MiMo-V2.5-Pro92.1%estimated ± 4.0 pp, low confidence
51MiniMax M391.0%estimated ± 4.0 pp, low confidence
52Qwen 3.6 Max (preview)91.0%estimated ± 4.0 pp, low confidence
53Claude Opus 4.689.1%measured
54GPT-6 Astra88.8%estimated ± 3.4 pp, low confidence
55Kimi K2.5 (Reasoning)88.8%estimated ± 4.0 pp, medium confidence
56Nemotron 3 Super 100B88.8%estimated ± 4.0 pp, medium confidence
57Sakana Fugu88.7%estimated ± 3.4 pp, low confidence
58Sakana Fugu-Ultra88.7%estimated ± 3.4 pp, low confidence
59GPT-5.6 Sol88.4%estimated ± 3.4 pp, low confidence
60Gemini 3.1 Pro88.3%estimated ± 3.4 pp, low confidence
61Claude Opus 4.7 (Adaptive)88.2%estimated ± 3.4 pp, low confidence
62Claude Opus 4.888.1%estimated ± 3.4 pp, low confidence
63GPT-5.588.1%estimated ± 3.4 pp, low confidence
64Kimi K388.0%estimated ± 3.4 pp, low confidence
65Step 5 Preview88.0%estimated ± 3.4 pp, low confidence
66Grok 4.1 Fast (Reasoning)88.0%estimated ± 4.0 pp, medium confidence
67Grok 4 Fast (Reasoning)88.0%estimated ± 4.0 pp, medium confidence
68Qwen3.6 Plus88.0%estimated ± 4.0 pp, medium confidence
69GPT-5.6 Terra87.8%estimated ± 3.4 pp, low confidence
70GPT-5.487.8%estimated ± 3.4 pp, low confidence
71Ornith-1.5-397B87.8%estimated ± 3.4 pp, low confidence
72Gemini 3.5 Flash87.8%estimated ± 3.4 pp, low confidence
73Qwen3.8 Max87.7%estimated ± 3.4 pp, low confidence
74Pareto 26.10 Preview87.7%estimated ± 3.4 pp, low confidence
75Qwen3.7 Max87.7%estimated ± 3.4 pp, low confidence
76GPT-5.6 Luna87.6%estimated ± 3.4 pp, low confidence
77Hy4 preview87.6%estimated ± 3.4 pp, low confidence
78Qwen3.8-Flash-Next87.4%estimated ± 3.4 pp, low confidence
79GLM-5.287.3%estimated ± 3.4 pp, low confidence
80Qwen3.8-Omni-Flash87.2%estimated ± 3.4 pp, low confidence
81DeepSeek V4.1 Flash87.2%estimated ± 3.4 pp, low confidence
82Kimi K2.587.1%measured
83Kimi K2.687.0%estimated ± 3.4 pp, low confidence
84Beam87.0%estimated ± 3.4 pp, low confidence
85Qwen3.7 Plus87.0%estimated ± 3.4 pp, low confidence
86DeepSeek V4 Pro 081386.9%estimated ± 3.4 pp, low confidence
87Gemini 2.5 Pro86.9%estimated ± 4.0 pp, medium confidence
88Muse Glimmer 30B86.9%estimated ± 4.0 pp, medium confidence
89Interfaze Beta86.8%estimated ± 3.4 pp, low confidence
90Inkling-Small86.7%estimated ± 3.4 pp, low confidence
91Muse Spark86.7%estimated ± 3.4 pp, low confidence
92Ornith-1.5-35B-A3B86.6%estimated ± 3.4 pp, medium confidence
93Qwen3.8-27B86.6%estimated ± 3.4 pp, medium confidence
94Solar Pro 486.5%estimated ± 3.4 pp, medium confidence
95Grok 4.2086.4%estimated ± 3.4 pp, medium confidence
96DeepSeek V4 Flash 073186.2%estimated ± 3.4 pp, medium confidence
97Inkling86.2%estimated ± 3.4 pp, medium confidence
98Hy3 Preview85.9%estimated ± 3.4 pp, medium confidence
99Nemotron 3 Ultra85.8%estimated ± 3.4 pp, medium confidence
100GLM-585.8%measured
101Step 3.7 Flash85.8%estimated ± 4.0 pp, medium confidence
102Ornith-1.5-9B85.6%estimated ± 3.4 pp, medium confidence
103Solar Open 285.6%estimated ± 3.4 pp, medium confidence
104GLM-5.185.6%estimated ± 3.4 pp, medium confidence
105Ternary Bonsai 2 27B85.4%estimated ± 3.4 pp, medium confidence
106A.X K285.3%estimated ± 3.4 pp, medium confidence
107Ling 3.0 Flash85.1%estimated ± 3.4 pp, medium confidence
108MAI-Thinking-184.8%estimated ± 3.4 pp, medium confidence
109Ling 3.0 Flash FP884.8%estimated ± 3.4 pp, medium confidence
110DeepSeek V3.1 (Reasoning)84.6%estimated ± 4.0 pp, medium confidence
111Grok 484.6%estimated ± 4.0 pp, medium confidence
112Ling 3.0 Flash VL84.6%estimated ± 4.0 pp, medium confidence
113o3-pro84.3%estimated ± 5.4 pp, low confidence
114K-EXAONE 2.084.1%estimated ± 3.4 pp, medium confidence
115GLM-4.783.5%estimated ± 4.0 pp, medium confidence
116Trinity-Large-Thinking83.4%measured
117Mercury 2.582.8%estimated ± 3.4 pp, medium confidence
118Gemma 4 12B82.8%estimated ± 3.4 pp, medium confidence
119Claude 4.1 Opus82.7%estimated ± 5.4 pp, low confidence
120DeepSeek-R182.4%estimated ± 4.0 pp, medium confidence
121Gemini 2.5 Flash82.4%estimated ± 4.0 pp, medium confidence
122Gemini 3 Flash82.4%estimated ± 4.0 pp, medium confidence
123Gemma 4 31B82.4%estimated ± 4.0 pp, medium confidence
124GPT-OSS 20B82.4%estimated ± 4.0 pp, medium confidence
125Nemotron 3.5 Lightning 30B A3B NVFP481.4%estimated ± 3.4 pp, medium confidence
126Claude 4 Sonnet81.3%estimated ± 4.0 pp, medium confidence
127GPT-OSS 120B81.3%estimated ± 4.0 pp, medium confidence
128K-Exaone81.3%estimated ± 4.0 pp, medium confidence
129MiMo-V2-Omni81.3%estimated ± 4.0 pp, medium confidence
130o381.3%estimated ± 4.0 pp, medium confidence
131Qwen3.6-27B81.3%estimated ± 4.0 pp, medium confidence
132MiniMax M2.780.8%measured
133Claude Sonnet 4.680.5%estimated ± 4.0 pp, medium confidence
134DeepSeek V3.280.5%estimated ± 4.0 pp, medium confidence
135Nemotron 3 Nano 30B80.5%estimated ± 4.0 pp, medium confidence
136Qwen3.5-27B80.5%estimated ± 4.0 pp, medium confidence
137Qwen3.5-35B-A3B80.5%estimated ± 4.0 pp, medium confidence
138Qwen3.5 397B80.5%estimated ± 4.0 pp, medium confidence
139Qwen3.5 397B (Reasoning)80.5%estimated ± 4.0 pp, medium confidence
140Nemotron 3 Nano Omni 30B A3B80.0%estimated ± 3.4 pp, medium confidence
141o3-mini79.9%estimated ± 5.4 pp, low confidence
142o1-pro79.9%estimated ± 5.4 pp, low confidence
143ZAYA1-8B79.4%estimated ± 3.4 pp, medium confidence
144Gemma 4 E4B79.4%estimated ± 4.0 pp, medium confidence
145GLM-5V-Turbo79.4%estimated ± 4.0 pp, medium confidence
146Qwen3.5-122B-A10B79.4%estimated ± 4.0 pp, medium confidence
147o1-preview79.4%estimated ± 5.4 pp, low confidence
148MiniCPM5-2B79.1%estimated ± 3.4 pp, medium confidence
149LongCat-Flash-Lite-Sparse78.7%estimated ± 3.4 pp, medium confidence
150Claude Opus 4.578.3%estimated ± 4.0 pp, low confidence
151Command A+78.3%estimated ± 4.0 pp, low confidence
152GLM-5-Turbo78.3%estimated ± 4.0 pp, low confidence
153Granite 4.2 30B78.3%estimated ± 4.0 pp, low confidence
154Granite 4.2 8B78.3%estimated ± 4.0 pp, low confidence
155MiMo-V2-Pro78.3%estimated ± 4.0 pp, low confidence
156Mistral Small 478.3%estimated ± 4.0 pp, low confidence
157Mistral Small 4 (Reasoning)78.3%estimated ± 4.0 pp, low confidence
158North Mini Code78.3%estimated ± 4.0 pp, low confidence
159o178.3%estimated ± 4.0 pp, low confidence
160Qwen3.6-35B-A3B78.3%estimated ± 4.0 pp, low confidence
161Sarvam 30B78.3%estimated ± 4.0 pp, low confidence
162Claude 3 Opus78.1%estimated ± 5.4 pp, low confidence
163DeepSeek R1 Distill Qwen 32B78.0%estimated ± 5.4 pp, low confidence
164Gemini 1.5 Pro77.7%estimated ± 5.4 pp, low confidence
165GPT-4 Turbo77.3%estimated ± 5.4 pp, low confidence
166Qwen2.5 Coder 32B Instruct77.2%estimated ± 5.4 pp, low confidence
167GPT-4o mini77.2%estimated ± 5.4 pp, low confidence
168Celeris-177.1%estimated ± 4.0 pp, low confidence
169Claude 3 Haiku77.1%estimated ± 4.0 pp, low confidence
170Claude 4.1 Opus Thinking77.1%estimated ± 4.0 pp, low confidence
171DeepSeek V377.1%estimated ± 4.0 pp, low confidence
172DeepSeek V3 032477.1%estimated ± 4.0 pp, low confidence
173DeepSeek V3.177.1%estimated ± 4.0 pp, low confidence
174Exaone 4.0 1.2B77.1%estimated ± 4.0 pp, low confidence
175Exaone 4.0 32B77.1%estimated ± 4.0 pp, low confidence
176Gemini 3.5 Flash-Lite77.1%estimated ± 4.0 pp, low confidence
177Gemma 3 27B77.1%estimated ± 4.0 pp, low confidence
178Gemma 4 26B A4B77.1%estimated ± 4.0 pp, low confidence
179Gemma 4 E2B77.1%estimated ± 4.0 pp, low confidence
180GLM-4.5-Air77.1%estimated ± 4.0 pp, low confidence
181GLM-4.677.1%estimated ± 4.0 pp, low confidence
182GPT-4.177.1%estimated ± 4.0 pp, low confidence
183GPT-4.1 mini77.1%estimated ± 4.0 pp, low confidence
184GPT-4.1 nano77.1%estimated ± 4.0 pp, low confidence
185GPT-4o77.1%estimated ± 4.0 pp, low confidence
186GPT-5 (medium)77.1%estimated ± 4.0 pp, low confidence
187Granite-4.0-350M77.1%estimated ± 4.0 pp, low confidence
188Granite-4.0-H-1B77.1%estimated ± 4.0 pp, low confidence
189Granite-4.0-H-350M77.1%estimated ± 4.0 pp, low confidence
190Granite 4.2 3B77.1%estimated ± 4.0 pp, low confidence
191Grok 4.1 Fast77.1%estimated ± 4.0 pp, low confidence
192Grok Code Fast 177.1%estimated ± 4.0 pp, low confidence
193Kimi K277.1%estimated ± 4.0 pp, low confidence
194LFM2.5-2.6B77.1%estimated ± 4.0 pp, low confidence
195LFM2.5-8B-A1B77.1%estimated ± 4.0 pp, low confidence
196LFM2.5-VL-1.6B-Extract77.1%estimated ± 4.0 pp, low confidence
197Ling 2.6 Flash77.1%estimated ± 4.0 pp, low confidence
198Ling 3.0 Tiny77.1%estimated ± 4.0 pp, low confidence
199Llama 3.1 405B77.1%estimated ± 4.0 pp, low confidence
200Llama 4 Maverick77.1%estimated ± 4.0 pp, low confidence
201Llama 4 Scout77.1%estimated ± 4.0 pp, low confidence
202MiMo-V2-Flash77.1%estimated ± 4.0 pp, low confidence
203Mistral Large 277.1%estimated ± 4.0 pp, low confidence
204Mistral Large 377.1%estimated ± 4.0 pp, low confidence
205Mistral Medium 377.1%estimated ± 4.0 pp, low confidence
206Mistral Medium 3.5 128B77.1%estimated ± 4.0 pp, low confidence
207Nemotron Ultra 253B77.1%estimated ± 4.0 pp, low confidence
208Nova Pro77.1%estimated ± 4.0 pp, low confidence
209Phi-477.1%estimated ± 4.0 pp, low confidence
210Qwen3 Max77.1%estimated ± 4.0 pp, low confidence
211Qwen3-Omni-30B-A3B-Instruct77.1%estimated ± 4.0 pp, low confidence
212Qwen3-Omni-30B-A3B-Thinking77.1%estimated ± 4.0 pp, low confidence
213Sarvam 105B77.1%estimated ± 4.0 pp, low confidence
214Solar Pro 277.1%estimated ± 4.0 pp, low confidence
215Solar Pro 377.1%estimated ± 4.0 pp, low confidence
216Ultravox v0.6 Llama 3.3 70B77.1%estimated ± 4.0 pp, low confidence
217Phi-4 Multimodal Instruct76.8%estimated ± 5.4 pp, low confidence
218Gemini 1.0 Pro76.5%estimated ± 5.4 pp, low confidence
219Trinity-Large-Preview75.2%measured
220Mellum2-12B-A2.5B-Thinking72.6%estimated ± 3.4 pp, low confidence
221ZAYA1-74B-Preview72.4%estimated ± 3.4 pp, low confidence
222LLaDA2.2-mini63.9%estimated ± 3.4 pp, low confidence
223Soofi S 30B-A3B63.1%estimated ± 3.4 pp, low confidence
224Mellum2-12B-A2.5B-Instruct61.2%estimated ± 3.4 pp, low confidence
225MiniCPM5-1B47.0%estimated ± 3.4 pp, low confidence
226LFM2.5-230M46.0%estimated ± 3.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General