benchgap
Coding

CursorBench 4.0 leaderboard

As of 2026-10-07, the highest measured score on CursorBench 4.0 is 57.8% by Claude Opus 5.5. 155 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.557.8%measured
2Claude Sonnet 5.555.5%measured
3Claude Fable 5.151.8%measured
4Claude Mythos 551.0%estimated ± 3.6 pp, high confidence
5MiMo-V2.6-Pro50.0%estimated ± 3.5 pp, high confidence
6Gemini 4 Argon49.9%estimated ± 2.9 pp, high confidence
7Claude Opus 546.6%measured
8Step 5 Preview46.5%estimated ± 3.5 pp, high confidence
9Grok 4.746.3%measured
10Sakana Fugu-Ultra46.0%estimated ± 3.6 pp, high confidence
11GPT-6 Sol44.1%estimated ± 3.5 pp, high confidence
12GPT-6 Astra42.5%estimated ± 1.6 pp, high confidence
13GLM-5.3-Flash42.0%estimated ± 2.9 pp, medium confidence
14Qwen3.8 Max42.0%estimated ± 2.9 pp, medium confidence
15Claude Fable 541.7%estimated ± 1.6 pp, high confidence
16GPT-5.6 Sol41.7%measured
17Muse Spark 1.341.6%measured
18Grok 4.641.4%measured
19GPT-5.6 Terra41.3%measured
20Kimi K341.3%estimated ± 1.6 pp, high confidence
21Gemini 3.7 Flash41.1%estimated ± 1.6 pp, high confidence
22Hy4 preview39.8%estimated ± 3.6 pp, high confidence
23Beam39.7%estimated ± 3.6 pp, high confidence
24Gemini 3.8 Flash39.6%measured
25Ornith-1.5-397B39.4%estimated ± 3.6 pp, high confidence
26GPT-5.539.2%estimated ± 1.6 pp, high confidence
27GLM-5.339.0%estimated ± 1.6 pp, high confidence
28Claude Haiku 5.539.0%estimated ± 3.5 pp, high confidence
29Claude Opus 4.838.2%estimated ± 1.6 pp, high confidence
30GPT-6 Luna38.1%estimated ± 3.5 pp, high confidence
31Qwen3.8-Omni-Flash38.0%estimated ± 3.6 pp, high confidence
32GPT-6.1 Sol37.3%estimated ± 3.5 pp, high confidence
33Mistral Large 437.3%estimated ± 3.5 pp, high confidence
34Claude Opus 4.7 (Adaptive)37.3%estimated ± 1.6 pp, high confidence
35Ornith-1.0-397B37.2%estimated ± 3.6 pp, medium confidence
36Ling 3.1 Flash37.1%estimated ± 3.5 pp, high confidence
37Qwen3.8-Flash-Next36.5%estimated ± 1.6 pp, high confidence
38dots3-note Preview36.2%estimated ± 3.6 pp, medium confidence
39GPT-5.6 Luna35.9%measured
40Grok 4.535.7%estimated ± 1.6 pp, high confidence
41Muse Spark 1.235.4%estimated ± 1.6 pp, high confidence
42Atria Dawn Preview35.2%estimated ± 3.6 pp, medium confidence
43Ornith-1.5-35B-A3B35.2%estimated ± 3.6 pp, medium confidence
44Claude Haiku 4.535.1%estimated ± 3.6 pp, low confidence
45Laguna S 2.135.0%estimated ± 3.6 pp, medium confidence
46Qwen3.8 Max Preview34.8%estimated ± 1.6 pp, high confidence
47Sakana Fugu34.7%estimated ± 3.6 pp, medium confidence
48Muse Spark 1.134.2%estimated ± 1.6 pp, medium confidence
49Claude Sonnet 534.1%measured
50GPT-5.433.9%estimated ± 1.6 pp, medium confidence
51Qwen 3.6 Max (preview)33.4%estimated ± 3.6 pp, medium confidence
52Claude Opus 4.533.3%estimated ± 3.6 pp, medium confidence
53GPT-5.3 Codex33.0%estimated ± 3.6 pp, medium confidence
54Gemini 3.5 Flash32.7%estimated ± 1.6 pp, medium confidence
55MiMo-V2.532.5%estimated ± 3.6 pp, medium confidence
56Claude Opus 4.732.5%estimated ± 4.4 pp, medium confidence
57DeepSeek V4.1 Flash32.4%estimated ± 3.5 pp, medium confidence
58GPT-5.232.1%estimated ± 3.6 pp, medium confidence
59GLM-531.7%estimated ± 3.6 pp, medium confidence
60Gemini 3.6 Flash31.7%estimated ± 1.6 pp, medium confidence
61DeepSeek V4 Flash 073131.5%estimated ± 1.6 pp, medium confidence
62DeepSeek V4 Pro 081331.2%estimated ± 1.6 pp, medium confidence
63Gemini 3.1 Pro31.2%estimated ± 1.6 pp, medium confidence
64MiMo-V2.6-Flash31.2%estimated ± 3.5 pp, medium confidence
65GLM-5.231.1%estimated ± 1.6 pp, medium confidence
66Claude Opus 4.630.4%estimated ± 3.6 pp, medium confidence
67Qwen3.8-27B30.4%estimated ± 1.6 pp, medium confidence
68MAI-Thinking-130.0%estimated ± 3.6 pp, medium confidence
69Claude Sonnet 4.629.5%estimated ± 4.4 pp, low confidence
70Grok 4.2029.2%estimated ± 3.6 pp, medium confidence
71Qwen3.5 397B28.5%estimated ± 3.6 pp, medium confidence
72Qwen3.7 Max28.1%estimated ± 1.6 pp, medium confidence
73Ornith-1.0-35B28.1%estimated ± 3.6 pp, medium confidence
74Gemini 3 Flash28.0%estimated ± 4.4 pp, low confidence
75Composer 2.527.7%measured
76Laguna M.127.2%estimated ± 3.6 pp, medium confidence
77GPT-5.2-Codex26.5%estimated ± 4.4 pp, low confidence
78Laguna XS 2.126.0%estimated ± 3.6 pp, medium confidence
79Ornith-1.5-9B25.9%estimated ± 3.6 pp, medium confidence
80Laguna XS.225.0%estimated ± 3.6 pp, medium confidence
81Kimi K2.624.2%estimated ± 1.6 pp, medium confidence
82Quasar 438B23.8%estimated ± 1.6 pp, medium confidence
83Apodex 1.123.4%estimated ± 1.6 pp, medium confidence
84Apodex 1.1 Mini23.4%estimated ± 1.6 pp, medium confidence
85Kimi K2.7 Code23.4%estimated ± 1.6 pp, medium confidence
86MiMo-V2.5-Pro22.9%estimated ± 1.6 pp, medium confidence
87Ornith-1.0-9B22.4%estimated ± 3.6 pp, medium confidence
88Qwen3.5 Flash22.2%estimated ± 4.4 pp, low confidence
89Hy321.8%estimated ± 1.6 pp, medium confidence
90Hy3 Preview21.8%estimated ± 1.6 pp, medium confidence
91Muse Spark21.7%estimated ± 1.6 pp, medium confidence
92MiniMax M321.7%estimated ± 1.6 pp, medium confidence
93Gemini 3.1 Flash-Lite21.4%estimated ± 4.4 pp, low confidence
94LongCat-Flash-Lite-Sparse20.6%estimated ± 3.6 pp, medium confidence
95GPT-5.4 mini19.8%estimated ± 1.6 pp, medium confidence
96GPT-5.4 nano19.8%estimated ± 1.6 pp, medium confidence
97Qwen3.7 Plus19.7%estimated ± 1.6 pp, medium confidence
98GLM-5.119.6%estimated ± 1.6 pp, medium confidence
99Qwen3.6 Plus18.8%estimated ± 1.6 pp, medium confidence
100Qwen3.6-27B18.3%estimated ± 1.6 pp, medium confidence
101Inkling-Small17.8%estimated ± 1.6 pp, medium confidence
102Solar Pro 417.6%estimated ± 3.5 pp, medium confidence
103MiniMax M2.717.6%estimated ± 1.6 pp, medium confidence
104Inkling17.2%estimated ± 1.6 pp, medium confidence
105Ling 3.0 Flash VL16.9%estimated ± 3.5 pp, medium confidence
106Ling 3.0 Flash16.4%estimated ± 1.6 pp, medium confidence
107Ling 3.0 Flash FP816.4%estimated ± 1.6 pp, medium confidence
108MiMo-V2-Flash15.9%estimated ± 1.6 pp, medium confidence
109GPT-5.115.7%estimated ± 1.6 pp, medium confidence
110Gemini 3.5 Flash-Lite15.6%estimated ± 1.6 pp, medium confidence
111Nemotron 3 Ultra15.6%estimated ± 1.6 pp, medium confidence
112Muse Glimmer 30B15.5%estimated ± 1.6 pp, medium confidence
113Mistral Medium 3.5 128B14.3%estimated ± 1.6 pp, medium confidence
114Kimi K2.514.3%estimated ± 1.6 pp, medium confidence
115Kimi K2.5 (Reasoning)14.3%estimated ± 1.6 pp, medium confidence
116Qwen3.5-122B-A10B13.7%estimated ± 1.6 pp, medium confidence
117GLM-4.713.5%estimated ± 1.6 pp, medium confidence
118K-EXAONE 2.013.2%estimated ± 3.5 pp, medium confidence
119Gemma 4 31B12.6%estimated ± 1.6 pp, medium confidence
120LLaDA2.2-flash12.6%estimated ± 3.6 pp, medium confidence
121Grok 4.312.1%estimated ± 1.6 pp, medium confidence
122Qwen3.6-35B-A3B11.9%estimated ± 1.6 pp, medium confidence
123A.X K211.7%estimated ± 3.5 pp, medium confidence
124o111.0%estimated ± 1.6 pp, medium confidence
125Step 3.7 Flash10.9%estimated ± 1.6 pp, medium confidence
126Gemma 4 26B A4B10.8%estimated ± 1.6 pp, medium confidence
127GPT-5 (high)10.2%estimated ± 1.6 pp, medium confidence
128Nemotron 3 Super 100B10.1%estimated ± 1.6 pp, medium confidence
129DeepSeek V3 03249.0%estimated ± 3.5 pp, medium confidence
130North Mini Code8.8%estimated ± 3.5 pp, medium confidence
131o1-preview8.7%estimated ± 1.6 pp, medium confidence
132Gemini 2.5 Pro8.5%estimated ± 1.6 pp, medium confidence
133Mercury 2.58.4%estimated ± 3.5 pp, medium confidence
134K-Exaone8.0%estimated ± 1.6 pp, medium confidence
135Gemma 4 12B7.7%estimated ± 1.6 pp, medium confidence
136Granite 4.2 30B7.6%estimated ± 3.5 pp, medium confidence
137GPT-OSS 120B7.5%estimated ± 1.6 pp, medium confidence
138Command A+6.6%estimated ± 1.6 pp, medium confidence
139Nemotron 3.5 Lightning 30B A3B NVFP46.3%estimated ± 1.6 pp, medium confidence
140Mistral Small 46.3%estimated ± 1.6 pp, medium confidence
141Mistral Small 4 (Reasoning)6.3%estimated ± 1.6 pp, medium confidence
142Trinity-Large-Preview6.0%estimated ± 1.6 pp, medium confidence
143Trinity-Large-Thinking6.0%estimated ± 1.6 pp, medium confidence
144Ling 2.6 Flash5.8%estimated ± 1.6 pp, medium confidence
145Gemini 1.5 Pro5.4%estimated ± 1.6 pp, medium confidence
146DeepSeek V35.2%estimated ± 1.6 pp, medium confidence
147Granite 4.2 8B5.0%estimated ± 1.6 pp, medium confidence
148GPT-4 Turbo4.8%estimated ± 1.6 pp, medium confidence
149GPT-OSS 20B4.6%estimated ± 1.6 pp, medium confidence
150GPT-4.1 mini4.4%estimated ± 1.6 pp, medium confidence
151Mistral Large 34.4%estimated ± 1.6 pp, medium confidence
152Claude 3 Opus4.2%estimated ± 1.6 pp, medium confidence
153Llama 4 Maverick3.4%estimated ± 1.6 pp, medium confidence
154Celeris-13.0%estimated ± 1.6 pp, medium confidence
155Nemotron 3 Nano 30B3.0%estimated ± 1.6 pp, medium confidence
156Nemotron 3 Nano Omni 30B A3B2.8%estimated ± 1.6 pp, medium confidence
157Ultravox v0.6 Llama 3.3 70B2.4%estimated ± 1.6 pp, medium confidence
158GPT-4o mini2.3%estimated ± 1.6 pp, medium confidence
159GPT-4.1 nano2.2%estimated ± 1.6 pp, medium confidence
160Gemma 3 27B2.0%estimated ± 1.6 pp, medium confidence
161Gemma 4 E4B1.8%estimated ± 1.6 pp, medium confidence
162Llama 4 Scout1.6%estimated ± 1.6 pp, medium confidence
163LFM2.5-2.6B1.5%estimated ± 1.6 pp, medium confidence
164Gemma 4 E2B1.4%estimated ± 1.6 pp, medium confidence
165MiniCPM5-2B1.0%estimated ± 3.5 pp, medium confidence
166Solar Pro 30.8%estimated ± 3.5 pp, medium confidence
167Granite 4.2 3B0.8%estimated ± 3.5 pp, medium confidence
168Ling 3.0 Tiny0.6%estimated ± 3.5 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General