benchgap
Knowledge & reasoning

C-Eval leaderboard

As of 2026-10-07, the highest measured score on C-Eval is 93.3% by Qwen3.6 Plus. 225 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 4.6100.0%estimated ± 0.7 pp, low confidence
2Claude Sonnet 4.6100.0%estimated ± 0.7 pp, low confidence
3Gemini 3 Pro Deep Think100.0%estimated ± 1.4 pp, low confidence
4Claude Opus 5.597.1%estimated ± 1.0 pp, low confidence
5Claude Sonnet 5.597.0%estimated ± 1.0 pp, low confidence
6Gemini 4 Argon96.8%estimated ± 1.0 pp, low confidence
7GPT-6.1 Sol96.7%estimated ± 1.0 pp, low confidence
8Claude Fable 596.6%estimated ± 1.0 pp, low confidence
9Muse Spark 1.396.4%estimated ± 1.0 pp, low confidence
10GPT-6 Sol96.4%estimated ± 1.0 pp, low confidence
11GPT-6 Astra96.4%estimated ± 0.9 pp, low confidence
12Grok 4.796.3%estimated ± 1.0 pp, low confidence
13MiMo-V2.6-Pro96.3%estimated ± 1.0 pp, low confidence
14Qwen3.8 Max Preview96.2%estimated ± 1.0 pp, low confidence
15GLM-5.396.1%estimated ± 1.0 pp, low confidence
16Sakana Fugu96.1%estimated ± 0.9 pp, low confidence
17Sakana Fugu-Ultra96.1%estimated ± 0.9 pp, low confidence
18Grok 4.696.1%estimated ± 1.0 pp, low confidence
19Claude Haiku 5.596.0%estimated ± 1.0 pp, low confidence
20GLM-5.3-Flash95.8%estimated ± 1.0 pp, low confidence
21Ling 3.1 Flash95.8%estimated ± 1.0 pp, low confidence
22Gemini 3.8 Flash95.8%estimated ± 1.0 pp, low confidence
23GPT-5.6 Sol95.6%estimated ± 0.9 pp, low confidence
24Muse Spark 1.295.6%estimated ± 1.0 pp, low confidence
25Gemini 3.7 Flash95.5%estimated ± 1.0 pp, low confidence
26Grok 4.595.5%estimated ± 1.0 pp, low confidence
27Mistral Large 495.5%estimated ± 1.0 pp, low confidence
28GPT-6 Luna95.4%estimated ± 1.0 pp, low confidence
29MiMo-V2.6-Flash95.4%estimated ± 1.0 pp, low confidence
30Gemini 3.6 Flash94.8%estimated ± 1.0 pp, low confidence
31GPT-5.6 Terra94.7%estimated ± 0.9 pp, low confidence
32GPT-5.3 Codex94.6%estimated ± 1.0 pp, low confidence
33Claude Opus 4.6 (Adaptive)94.5%estimated ± 1.0 pp, low confidence
34GPT-5.294.5%estimated ± 0.9 pp, low confidence
35GPT-5.6 Luna94.4%estimated ± 0.9 pp, low confidence
36Claude Opus 4.794.3%estimated ± 1.0 pp, low confidence
37Gemini 3.1 Pro94.1%estimated ± 1.0 pp, low confidence
38Qwen 3.6 Max (preview)94.0%estimated ± 0.7 pp, low confidence
39MiniMax M394.0%estimated ± 1.0 pp, low confidence
40Claude Opus 4.5 Thinking93.9%estimated ± 1.0 pp, low confidence
41Qwen3.7 Max93.9%estimated ± 0.7 pp, low confidence
42MiMo-V2-Pro93.8%estimated ± 1.0 pp, low confidence
43GPT-5.2-Codex93.8%estimated ± 1.0 pp, low confidence
44Solar Pro 493.7%estimated ± 1.0 pp, low confidence
45Gemini 3 Pro93.7%estimated ± 1.0 pp, low confidence
46Quasar 438B93.4%estimated ± 1.0 pp, medium confidence
47GLM-5-Turbo93.4%estimated ± 1.0 pp, medium confidence
48Apodex 1.1 Mini93.3%estimated ± 1.0 pp, medium confidence
49Qwen3.6 Plus93.3%measured
50Claude Fable 5.193.3%estimated ± 0.8 pp, low confidence
51Claude Opus 593.3%estimated ± 0.8 pp, low confidence
52Claude Mythos 593.3%estimated ± 0.8 pp, low confidence
53Muse Spark 1.193.3%estimated ± 0.8 pp, low confidence
54GPT-5.4 Pro93.3%estimated ± 0.8 pp, low confidence
55Claude Opus 4.893.3%estimated ± 0.8 pp, low confidence
56Claude Sonnet 593.3%estimated ± 0.8 pp, low confidence
57GPT-5.5 Pro93.3%estimated ± 0.8 pp, low confidence
58Apodex 1.193.3%estimated ± 0.8 pp, low confidence
59Kimi K393.3%estimated ± 0.8 pp, low confidence
60Hy4 preview93.3%estimated ± 0.8 pp, low confidence
61Claude Opus 4.7 (Adaptive)93.3%estimated ± 0.8 pp, low confidence
62GLM-5.293.3%estimated ± 0.8 pp, low confidence
63dots3-note Preview93.3%estimated ± 0.8 pp, low confidence
64GLM-5.193.3%estimated ± 0.8 pp, low confidence
65GPT-5.593.3%estimated ± 0.8 pp, low confidence
66GPT-5.493.3%estimated ± 0.8 pp, low confidence
67Muse Spark93.3%estimated ± 0.8 pp, low confidence
68MiMo-V2.5-Pro93.3%estimated ± 0.8 pp, low confidence
69Inkling-Small93.3%estimated ± 0.8 pp, low confidence
70Agents-A193.3%estimated ± 0.8 pp, low confidence
71Step 5 Preview93.3%estimated ± 0.8 pp, low confidence
72Inkling93.3%estimated ± 0.8 pp, low confidence
73Ornith-1.5-397B93.3%estimated ± 0.8 pp, low confidence
74Qwen3.8 Max93.2%estimated ± 0.8 pp, low confidence
75DeepSeek V4 Pro 081393.2%estimated ± 0.8 pp, low confidence
76GPT-5.4 mini93.2%estimated ± 0.8 pp, low confidence
77Gemini 3.5 Flash93.2%estimated ± 0.8 pp, low confidence
78GPT-5.4 nano93.2%estimated ± 0.8 pp, low confidence
79DeepSeek V4.1 Flash93.2%estimated ± 0.8 pp, low confidence
80Qwen3.8-Omni-Flash93.2%estimated ± 0.8 pp, low confidence
81Kimi K2.7 Code93.1%estimated ± 1.0 pp, medium confidence
82Qwen3.8-Flash-Next93.1%estimated ± 0.8 pp, low confidence
83Grok 4.393.1%estimated ± 0.8 pp, low confidence
84DeepSeek V4 Flash 073193.1%estimated ± 0.8 pp, low confidence
85Kimi K2.693.1%estimated ± 0.8 pp, low confidence
86Interfaze Beta93.1%estimated ± 0.9 pp, medium confidence
87Qwen3.7 Plus93.1%estimated ± 0.7 pp, medium confidence
88Hy393.0%estimated ± 1.0 pp, medium confidence
89Qwen3.5 397B93.0%measured
90Qwen3.8-27B92.9%estimated ± 0.8 pp, medium confidence
91GPT-5.192.9%estimated ± 1.0 pp, medium confidence
92Seed 2.1 Pro92.8%estimated ± 0.7 pp, medium confidence
93Ling 3.0 Flash VL92.8%estimated ± 1.0 pp, medium confidence
94MiMo-V2-Omni92.6%estimated ± 1.0 pp, medium confidence
95A.X K292.6%estimated ± 0.8 pp, medium confidence
96GPT-5.1-Codex92.5%estimated ± 1.0 pp, medium confidence
97GPT-5.1-Codex-Max92.5%estimated ± 1.0 pp, medium confidence
98GLM-5V-Turbo92.5%estimated ± 1.0 pp, medium confidence
99Nemotron 3 Ultra92.4%estimated ± 0.8 pp, medium confidence
100Gemma 4 31B92.3%estimated ± 0.8 pp, medium confidence
101GPT-5 (high)92.3%estimated ± 1.0 pp, medium confidence
102GPT-5 (medium)92.3%estimated ± 1.0 pp, medium confidence
103Claude 4.1 Opus Thinking92.3%estimated ± 1.0 pp, medium confidence
104MiniMax M2.792.2%estimated ± 1.0 pp, medium confidence
105Kimi K2.592.2%estimated ± 0.7 pp, medium confidence
106Claude Opus 4.592.2%measured
107Command A+92.2%estimated ± 1.0 pp, medium confidence
108Grok 492.1%estimated ± 1.0 pp, medium confidence
109Ornith-1.5-35B-A3B92.1%estimated ± 0.8 pp, medium confidence
110Hy3 Preview92.1%estimated ± 0.8 pp, medium confidence
111Gemini 3.5 Flash-Lite92.0%estimated ± 1.0 pp, medium confidence
112o3-pro91.9%estimated ± 1.0 pp, medium confidence
113GLM-4.791.9%estimated ± 0.8 pp, medium confidence
114Kimi K2.5 (Reasoning)91.8%estimated ± 0.9 pp, medium confidence
115Qwen3.5 397B (Reasoning)91.8%estimated ± 1.0 pp, medium confidence
116Seed 2.1 Turbo91.5%estimated ± 0.7 pp, medium confidence
117Qwen3.6-27B91.4%measured
118Qwen3.5-122B-A10B91.4%estimated ± 0.7 pp, medium confidence
119Grok 4.1 Fast (Reasoning)91.3%estimated ± 1.0 pp, medium confidence
120o391.3%estimated ± 1.0 pp, medium confidence
121GLM-591.3%estimated ± 0.7 pp, medium confidence
122Solar Open 291.1%estimated ± 1.0 pp, medium confidence
123Step 3.7 Flash91.0%estimated ± 1.0 pp, medium confidence
124Ling 3.0 Flash90.9%estimated ± 0.8 pp, medium confidence
125Ternary Bonsai 2 27B90.8%estimated ± 0.9 pp, low confidence
126Qwen3.5-27B90.7%estimated ± 0.7 pp, medium confidence
127Claude 4.1 Opus90.5%estimated ± 1.0 pp, medium confidence
128Grok 4 Fast (Reasoning)90.2%estimated ± 1.0 pp, low confidence
129Gemini 3 Flash90.2%estimated ± 1.0 pp, low confidence
130Qwen3.6-35B-A3B90.0%measured
131Muse Glimmer 30B90.0%estimated ± 1.0 pp, low confidence
132MAI-Thinking-189.8%estimated ± 0.9 pp, low confidence
133Qwen3.5-35B-A3B89.8%estimated ± 0.7 pp, low confidence
134Ling 3.0 Flash FP889.7%estimated ± 0.9 pp, low confidence
135MiMo-V2-Flash89.5%estimated ± 0.9 pp, low confidence
136Claude 4 Sonnet89.5%estimated ± 1.0 pp, low confidence
137Qwen3 235B 250789.4%estimated ± 0.7 pp, low confidence
138Claude Sonnet 4.589.4%estimated ± 0.9 pp, low confidence
139DeepSeek V3.289.1%estimated ± 1.0 pp, low confidence
140Qwen3 Max88.9%estimated ± 1.0 pp, low confidence
141Ornith-1.5-9B88.7%estimated ± 0.8 pp, low confidence
142GLM-4.688.4%estimated ± 1.0 pp, low confidence
143K-Exaone88.0%estimated ± 1.0 pp, low confidence
144Mistral Medium 3.5 128B87.8%estimated ± 1.0 pp, low confidence
145Grok Code Fast 187.7%estimated ± 1.0 pp, low confidence
146DeepSeek V3.187.5%estimated ± 1.0 pp, low confidence
147DeepSeek V3.1 (Reasoning)87.3%estimated ± 1.0 pp, low confidence
148DeepSeek-R187.0%estimated ± 1.0 pp, low confidence
149o1-pro86.7%estimated ± 0.9 pp, low confidence
150Nemotron 3 Super 100B86.7%estimated ± 1.0 pp, low confidence
151Kimi K286.6%estimated ± 1.0 pp, low confidence
152Gemma 4 12B86.6%estimated ± 0.9 pp, low confidence
153Gemini 2.5 Pro86.4%estimated ± 0.8 pp, low confidence
154Mercury 2.586.2%estimated ± 1.0 pp, low confidence
155LongCat-Flash-Lite-Sparse85.8%measured
156o3-mini85.6%estimated ± 0.9 pp, low confidence
157GPT-OSS 120B85.4%estimated ± 1.0 pp, low confidence
158K-EXAONE 2.085.2%estimated ± 0.8 pp, low confidence
159o1-preview85.2%estimated ± 1.0 pp, low confidence
160Grok 4.1 Fast85.1%estimated ± 1.0 pp, low confidence
161Mistral Small 485.1%estimated ± 1.0 pp, low confidence
162Mistral Small 4 (Reasoning)85.1%estimated ± 1.0 pp, low confidence
163GLM-4.5-Air84.9%estimated ± 1.0 pp, low confidence
164Ling 3.0 Tiny84.8%estimated ± 1.0 pp, low confidence
165o184.6%estimated ± 0.9 pp, low confidence
166Nemotron 3.5 Lightning 30B A3B NVFP484.5%estimated ± 0.9 pp, low confidence
167Trinity-Large-Preview84.5%estimated ± 1.0 pp, low confidence
168Trinity-Large-Thinking84.5%estimated ± 1.0 pp, low confidence
169Llama 4 Maverick83.4%estimated ± 1.0 pp, low confidence
170North Mini Code83.3%estimated ± 1.0 pp, low confidence
171Gemini 2.5 Flash83.2%estimated ± 1.0 pp, low confidence
172DeepSeek V3 032483.0%estimated ± 1.0 pp, low confidence
173Mistral Large 382.3%estimated ± 1.0 pp, low confidence
174Nemotron 3 Nano Omni 30B A3B82.3%estimated ± 0.9 pp, low confidence
175Gemma 4 26B A4B82.1%estimated ± 0.8 pp, low confidence
176Mistral Medium 382.0%estimated ± 1.0 pp, low confidence
177GPT-OSS 20B81.8%estimated ± 1.0 pp, low confidence
178Nemotron 3 Nano 30B81.7%estimated ± 1.0 pp, low confidence
179Sarvam 105B81.5%estimated ± 1.0 pp, low confidence
180ZAYA1-8B81.5%estimated ± 0.9 pp, low confidence
181Claude 3 Opus81.4%estimated ± 1.0 pp, low confidence
182GPT-4o80.9%estimated ± 1.0 pp, low confidence
183LFM2.5-2.6B80.8%estimated ± 1.0 pp, low confidence
184DeepSeek R1 Distill Qwen 32B80.8%estimated ± 1.0 pp, low confidence
185Llama 4 Scout80.2%estimated ± 1.0 pp, low confidence
186Gemini 1.5 Pro79.8%estimated ± 1.0 pp, low confidence
187Solar Pro 379.6%estimated ± 1.0 pp, low confidence
188Qwen3-Omni-30B-A3B-Thinking79.5%estimated ± 1.0 pp, low confidence
189Ultravox v0.6 Llama 3.3 70B79.3%estimated ± 1.0 pp, low confidence
190Mistral Large 279.1%estimated ± 1.0 pp, low confidence
191Nemotron Ultra 253B79.0%estimated ± 1.0 pp, low confidence
192Llama 3.1 405B78.5%estimated ± 1.0 pp, low confidence
193LFM2.5-8B-A1B78.3%estimated ± 1.0 pp, low confidence
194Granite 4.2 30B78.3%estimated ± 0.9 pp, low confidence
195GPT-4.178.2%estimated ± 0.9 pp, low confidence
196GPT-4 Turbo77.9%estimated ± 1.0 pp, low confidence
197Solar Pro 277.8%estimated ± 1.0 pp, low confidence
198Nova Pro77.7%estimated ± 1.0 pp, low confidence
199Qwen2.5 Coder 32B Instruct77.1%estimated ± 1.0 pp, low confidence
200GPT-4o mini76.9%estimated ± 1.0 pp, low confidence
201GPT-4.1 mini76.7%estimated ± 0.9 pp, low confidence
202Granite 4.2 8B76.6%estimated ± 0.9 pp, low confidence
203Sarvam 30B76.6%estimated ± 1.0 pp, low confidence
204MiniCPM5-2B76.4%estimated ± 0.7 pp, low confidence
205Celeris-176.0%estimated ± 1.0 pp, low confidence
206Exaone 4.0 32B75.9%estimated ± 1.0 pp, low confidence
207Qwen3-Omni-30B-A3B-Instruct75.0%estimated ± 1.0 pp, low confidence
208Phi-474.7%estimated ± 1.0 pp, low confidence
209Phi-4 Multimodal Instruct74.3%estimated ± 1.0 pp, low confidence
210Claude 3 Haiku73.5%estimated ± 1.0 pp, low confidence
211Claude 3.5 Sonnet73.0%estimated ± 0.9 pp, low confidence
212DeepSeek V372.8%estimated ± 0.9 pp, low confidence
213Ling 2.6 Flash72.7%estimated ± 0.9 pp, low confidence
214Gemini 1.0 Pro72.6%estimated ± 1.0 pp, low confidence
215Gemma 4 E4B72.4%estimated ± 0.9 pp, low confidence
216Exaone 4.0 1.2B72.2%estimated ± 1.0 pp, low confidence
217Granite-4.0-H-1B72.1%estimated ± 1.0 pp, low confidence
218Mellum2-12B-A2.5B-Thinking71.6%estimated ± 0.9 pp, low confidence
219ZAYA1-74B-Preview71.4%estimated ± 0.9 pp, low confidence
220Gemma 3 27B70.7%estimated ± 1.0 pp, low confidence
221Granite-4.0-350M70.6%estimated ± 1.0 pp, low confidence
222Granite-4.0-H-350M70.6%estimated ± 1.0 pp, low confidence
223LFM2.5-VL-1.6B-Extract70.6%estimated ± 1.0 pp, low confidence
224Granite 4.2 3B69.4%estimated ± 0.9 pp, low confidence
225GPT-4.1 nano65.5%estimated ± 0.9 pp, low confidence
226Gemma 4 E2B59.2%estimated ± 0.9 pp, low confidence
227Soofi S 30B-A3B59.2%estimated ± 0.9 pp, low confidence
228MiniCPM5-1B58.0%estimated ± 0.7 pp, low confidence
229Mellum2-12B-A2.5B-Instruct56.8%estimated ± 0.9 pp, low confidence
230LFM2.5-VL-450M39.8%estimated ± 0.9 pp, low confidence
231LFM2.5-230M39.5%estimated ± 0.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General