benchgap
Coding

CursorBench 3.1 leaderboard

As of 2026-10-07, the highest measured score on CursorBench 3.1 is 70.6% by Claude Fable 5. 161 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.585.7%estimated ± 3.4 pp, low confidence
2Claude Fable 5.176.3%estimated ± 3.4 pp, low confidence
3Gemini 4 Argon73.1%estimated ± 3.4 pp, low confidence
4Claude Sonnet 5.571.1%estimated ± 3.4 pp, medium confidence
5MiMo-V2.6-Pro70.9%estimated ± 3.4 pp, medium confidence
6Claude Fable 570.6%measured
7Kimi K367.4%estimated ± 3.4 pp, medium confidence
8Claude Mythos 566.9%estimated ± 5.6 pp, low confidence
9GLM-5.366.2%estimated ± 3.4 pp, medium confidence
10Step 5 Preview65.9%estimated ± 3.4 pp, medium confidence
11Muse Spark 1.165.7%estimated ± 3.4 pp, medium confidence
12Muse Spark 1.365.7%estimated ± 3.4 pp, medium confidence
13Gemini 3.1 Pro65.5%estimated ± 3.4 pp, medium confidence
14Composer 2.563.2%measured
15Sakana Fugu-Ultra63.2%estimated ± 5.6 pp, low confidence
16GPT-6 Sol62.7%estimated ± 3.4 pp, medium confidence
17Grok 4.762.2%estimated ± 3.4 pp, medium confidence
18Muse Spark 1.262.2%estimated ± 3.4 pp, medium confidence
19Gemini 3.7 Flash61.8%estimated ± 3.4 pp, medium confidence
20GPT-5.6 Sol61.5%estimated ± 3.4 pp, medium confidence
21Gemini 3.8 Flash60.3%estimated ± 3.4 pp, medium confidence
22GPT-6 Astra60.0%estimated ± 3.4 pp, medium confidence
23Grok 4.660.0%estimated ± 3.4 pp, medium confidence
24Claude Opus 559.8%estimated ± 3.4 pp, medium confidence
25Qwen3.8 Max59.6%estimated ± 5.6 pp, low confidence
26GPT-5.559.2%measured
27Claude Opus 4.858.4%measured
28Hy4 preview58.3%estimated ± 5.6 pp, low confidence
29Beam58.2%estimated ± 5.6 pp, low confidence
30Ornith-1.5-397B57.9%estimated ± 5.6 pp, low confidence
31Qwen3.8-Omni-Flash56.8%estimated ± 5.6 pp, low confidence
32Claude Haiku 5.556.3%estimated ± 3.4 pp, medium confidence
33GPT-5.6 Terra56.3%estimated ± 3.4 pp, medium confidence
34Grok 4.556.3%estimated ± 3.4 pp, medium confidence
35Claude Opus 4.756.2%estimated ± 5.8 pp, low confidence
36Ornith-1.0-397B56.1%estimated ± 5.6 pp, low confidence
37GPT-6 Luna55.3%estimated ± 3.4 pp, medium confidence
38dots3-note Preview55.3%estimated ± 5.6 pp, low confidence
39Claude Opus 4.7 (Adaptive)55.2%estimated ± 4.8 pp, medium confidence
40Claude Sonnet 554.6%estimated ± 3.4 pp, medium confidence
41Atria Dawn Preview54.4%estimated ± 5.6 pp, low confidence
42Ornith-1.5-35B-A3B54.4%estimated ± 5.6 pp, low confidence
43GPT-6.1 Sol54.4%estimated ± 3.4 pp, medium confidence
44Mistral Large 454.4%estimated ± 3.4 pp, medium confidence
45Laguna S 2.154.2%estimated ± 5.6 pp, low confidence
46Ling 3.1 Flash54.1%estimated ± 3.4 pp, medium confidence
47Sakana Fugu54.0%estimated ± 5.6 pp, low confidence
48GPT-5.6 Luna52.9%estimated ± 3.4 pp, medium confidence
49Qwen 3.6 Max (preview)52.8%estimated ± 5.6 pp, low confidence
50Claude Opus 4.552.7%estimated ± 5.6 pp, low confidence
51GPT-5.3 Codex52.5%estimated ± 5.6 pp, low confidence
52Gemini 3.6 Flash52.4%estimated ± 3.4 pp, medium confidence
53MiMo-V2.552.0%estimated ± 5.6 pp, low confidence
54GPT-5.251.7%estimated ± 5.6 pp, low confidence
55GLM-551.3%estimated ± 5.6 pp, low confidence
56GPT-5.450.2%estimated ± 4.8 pp, medium confidence
57Claude Opus 4.650.1%estimated ± 5.6 pp, low confidence
58Gemini 3.5 Flash49.8%measured
59MAI-Thinking-149.7%estimated ± 5.6 pp, low confidence
60GPT-5.4 mini49.2%estimated ± 3.4 pp, medium confidence
61Qwen3.8 Max Preview49.2%estimated ± 3.4 pp, medium confidence
62Grok 4.2049.0%estimated ± 5.6 pp, low confidence
63Gemini 3 Flash49.0%estimated ± 5.8 pp, low confidence
64Claude Sonnet 4.648.8%measured
65DeepSeek V4.1 Flash48.7%estimated ± 3.4 pp, medium confidence
66Qwen3.5 397B48.4%estimated ± 5.6 pp, low confidence
67Ornith-1.0-35B48.0%estimated ± 5.6 pp, low confidence
68GLM-5.3-Flash47.9%estimated ± 3.4 pp, medium confidence
69Muse Spark47.9%estimated ± 4.8 pp, low confidence
70Qwen3.6 Plus47.9%estimated ± 4.8 pp, low confidence
71GPT-5.147.9%estimated ± 4.8 pp, low confidence
72MiMo-V2-Flash47.9%estimated ± 4.8 pp, low confidence
73Claude 3 Opus47.9%estimated ± 4.8 pp, low confidence
74Gemini 1.5 Pro47.9%estimated ± 4.8 pp, low confidence
75Gemma 4 12B47.9%estimated ± 4.8 pp, low confidence
76Gemma 4 E2B47.9%estimated ± 4.8 pp, low confidence
77GLM-4.747.9%estimated ± 4.8 pp, low confidence
78GPT-4.1 mini47.9%estimated ± 4.8 pp, low confidence
79GPT-4.1 nano47.9%estimated ± 4.8 pp, low confidence
80GPT-4 Turbo47.9%estimated ± 4.8 pp, low confidence
81GPT-4o mini47.9%estimated ± 4.8 pp, low confidence
82GPT-5 (high)47.9%estimated ± 4.8 pp, low confidence
83K-Exaone47.9%estimated ± 4.8 pp, low confidence
84Kimi K2.547.9%estimated ± 4.8 pp, low confidence
85Kimi K2.5 (Reasoning)47.9%estimated ± 4.8 pp, low confidence
86Ling 2.6 Flash47.9%estimated ± 4.8 pp, low confidence
87Nemotron 3 Nano Omni 30B A3B47.9%estimated ± 4.8 pp, low confidence
88o147.9%estimated ± 4.8 pp, low confidence
89o1-preview47.9%estimated ± 4.8 pp, low confidence
90Ultravox v0.6 Llama 3.3 70B47.9%estimated ± 4.8 pp, low confidence
91Kimi K2.647.6%measured
92MiMo-V2.6-Flash47.2%estimated ± 3.4 pp, low confidence
93Laguna M.147.1%estimated ± 5.6 pp, low confidence
94GLM-5.247.0%estimated ± 3.4 pp, low confidence
95DeepSeek V4 Pro 081346.5%estimated ± 3.4 pp, low confidence
96GPT-5.2-Codex46.3%estimated ± 5.8 pp, low confidence
97Laguna XS 2.145.9%estimated ± 5.6 pp, low confidence
98Ornith-1.5-9B45.9%estimated ± 5.6 pp, low confidence
99MiMo-V2.5-Pro45.5%estimated ± 3.4 pp, low confidence
100Qwen3.8-Flash-Next45.5%estimated ± 3.4 pp, low confidence
101Laguna XS.245.0%estimated ± 5.6 pp, low confidence
102DeepSeek V4 Flash 073144.7%estimated ± 3.4 pp, low confidence
103MiniMax M2.744.3%estimated ± 3.4 pp, low confidence
104Inkling-Small43.3%estimated ± 3.4 pp, low confidence
105Qwen3.7 Max42.8%estimated ± 3.4 pp, low confidence
106Ornith-1.0-9B42.4%estimated ± 5.6 pp, low confidence
107LongCat-Flash-Lite-Sparse40.6%estimated ± 5.6 pp, low confidence
108Hy340.6%estimated ± 3.4 pp, low confidence
109Hy3 Preview40.6%estimated ± 3.4 pp, low confidence
110Claude Haiku 4.540.4%estimated ± 5.8 pp, low confidence
111Grok 4.339.8%estimated ± 3.4 pp, low confidence
112Quasar 438B39.3%estimated ± 3.4 pp, low confidence
113Kimi K2.7 Code38.6%estimated ± 3.4 pp, low confidence
114Qwen3.5 Flash38.1%estimated ± 5.8 pp, low confidence
115GPT-5.4 nano37.1%estimated ± 3.4 pp, low confidence
116MiniMax M336.9%estimated ± 3.4 pp, low confidence
117Inkling36.6%estimated ± 3.4 pp, low confidence
118Gemini 3.1 Flash-Lite36.5%estimated ± 5.8 pp, low confidence
119Qwen3.8-27B35.6%estimated ± 3.4 pp, low confidence
120Gemini 2.5 Pro34.9%estimated ± 3.4 pp, low confidence
121Qwen3.7 Plus34.4%estimated ± 3.4 pp, low confidence
122Apodex 1.132.9%estimated ± 3.4 pp, low confidence
123Apodex 1.1 Mini32.9%estimated ± 3.4 pp, low confidence
124Gemma 4 31B32.9%estimated ± 3.4 pp, low confidence
125LLaDA2.2-flash31.7%estimated ± 5.6 pp, low confidence
126Muse Glimmer 30B31.4%estimated ± 3.4 pp, low confidence
127GLM-5.131.2%estimated ± 3.4 pp, low confidence
128Solar Pro 430.7%estimated ± 3.4 pp, low confidence
129Ling 3.0 Flash VL29.7%estimated ± 3.4 pp, low confidence
130Step 3.7 Flash29.0%estimated ± 3.4 pp, low confidence
131Qwen3.6-27B26.3%estimated ± 3.4 pp, low confidence
132K-EXAONE 2.024.3%estimated ± 3.4 pp, low confidence
133Ling 3.0 Flash24.3%estimated ± 3.4 pp, low confidence
134Ling 3.0 Flash FP824.3%estimated ± 3.4 pp, low confidence
135Gemini 3.5 Flash-Lite22.6%estimated ± 3.4 pp, low confidence
136A.X K221.8%estimated ± 3.4 pp, low confidence
137Trinity-Large-Preview20.8%estimated ± 3.4 pp, low confidence
138Trinity-Large-Thinking20.8%estimated ± 3.4 pp, low confidence
139Nemotron 3 Ultra20.1%estimated ± 3.4 pp, low confidence
140Mistral Medium 3.5 128B19.8%estimated ± 3.4 pp, low confidence
141Gemma 4 26B A4B19.3%estimated ± 3.4 pp, low confidence
142Qwen3.5-122B-A10B18.6%estimated ± 3.4 pp, low confidence
143DeepSeek V3 032416.9%estimated ± 3.4 pp, low confidence
144GPT-OSS 20B16.6%estimated ± 3.4 pp, low confidence
145Mistral Small 416.4%estimated ± 3.4 pp, low confidence
146Mistral Small 4 (Reasoning)16.4%estimated ± 3.4 pp, low confidence
147North Mini Code16.4%estimated ± 3.4 pp, low confidence
148Command A+15.6%estimated ± 3.4 pp, low confidence
149Mercury 2.515.6%estimated ± 3.4 pp, low confidence
150Granite 4.2 30B13.9%estimated ± 3.4 pp, low confidence
151Mistral Large 311.0%estimated ± 3.4 pp, low confidence
152Qwen3.6-35B-A3B11.0%estimated ± 3.4 pp, low confidence
153Nemotron 3 Super 100B10.0%estimated ± 3.4 pp, low confidence
154DeepSeek V39.0%estimated ± 3.4 pp, low confidence
155GPT-OSS 120B4.6%estimated ± 3.4 pp, low confidence
156Celeris-10.0%estimated ± 3.4 pp, low confidence
157Gemma 3 27B0.0%estimated ± 3.4 pp, low confidence
158Gemma 4 E4B0.0%estimated ± 3.4 pp, low confidence
159Granite 4.2 3B0.0%estimated ± 3.4 pp, low confidence
160Granite 4.2 8B0.0%estimated ± 3.4 pp, low confidence
161LFM2.5-2.6B0.0%estimated ± 3.4 pp, low confidence
162Ling 3.0 Tiny0.0%estimated ± 3.4 pp, low confidence
163Llama 4 Maverick0.0%estimated ± 3.4 pp, low confidence
164Llama 4 Scout0.0%estimated ± 3.4 pp, low confidence
165MiniCPM5-2B0.0%estimated ± 3.4 pp, low confidence
166Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 3.4 pp, low confidence
167Nemotron 3 Nano 30B0.0%estimated ± 3.4 pp, low confidence
168Solar Pro 30.0%estimated ± 3.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General