benchgap
Coding

CursorBench 3.2 leaderboard

As of 2026-10-07, the highest measured score on CursorBench 3.2 is 73.4% by Claude Fable 5.1. 160 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.173.4%measured
2Gemini 4 Argon72.4%estimated ± 4.1 pp, low confidence
3Claude Mythos 571.4%estimated ± 3.6 pp, high confidence
4Grok 4.670.8%measured
5Grok 4.770.6%estimated ± 2.2 pp, high confidence
6Claude Fable 570.5%measured
7Claude Opus 5.570.5%estimated ± 1.6 pp, low confidence
8Claude Opus 570.0%measured
9GPT-6 Astra69.5%estimated ± 1.6 pp, medium confidence
10Gemini 3.8 Flash69.2%measured
11Sakana Fugu-Ultra69.1%estimated ± 3.6 pp, high confidence
12Ember-168.2%estimated ± 6.4 pp, low confidence
13Pareto 26.968.1%estimated ± 6.4 pp, low confidence
14DeepSeek V4 Flash 073167.8%estimated ± 3.3 pp, medium confidence
15Muse Spark 1.367.6%estimated ± 2.2 pp, high confidence
16Pareto 26.10 Preview67.6%estimated ± 6.4 pp, low confidence
17GPT-5.6 Sol67.2%measured
18Muse Spark 1.267.1%estimated ± 3.3 pp, medium confidence
19GLM-5.3-Flash67.0%estimated ± 4.1 pp, low confidence
20Grok 4.566.7%measured
21SWE-266.7%estimated ± 1.6 pp, medium confidence
22GPT-6 Sol65.8%estimated ± 5.3 pp, medium confidence
23Step 5 Preview65.3%estimated ± 3.6 pp, medium confidence
24Qwen3.8-27B64.9%estimated ± 3.3 pp, medium confidence
25GPT-5.6 Terra64.9%measured
26Qwen3.8 Max64.2%estimated ± 3.3 pp, medium confidence
27Hy4 preview64.2%estimated ± 3.6 pp, high confidence
28Beam64.0%estimated ± 3.6 pp, high confidence
29Ornith-1.5-397B63.7%estimated ± 3.6 pp, high confidence
30Claude Haiku 5.563.4%estimated ± 1.6 pp, medium confidence
31Claude Sonnet 5.563.2%estimated ± 1.6 pp, medium confidence
32Claude Opus 4.7 (Adaptive)62.9%estimated ± 3.6 pp, high confidence
33GLM-5.362.7%estimated ± 3.3 pp, medium confidence
34Claude Opus 4.862.3%measured
35Qwen3.8-Omni-Flash62.0%estimated ± 3.6 pp, high confidence
36Claude Haiku 4.561.5%estimated ± 3.3 pp, medium confidence
37Claude Sonnet 561.5%measured
38Qwen3.8-Flash-Next61.1%estimated ± 3.6 pp, high confidence
39GPT-5.6 Luna61.1%measured
40GPT-6 Luna61.0%estimated ± 5.3 pp, medium confidence
41Kimi K360.8%measured
42Ornith-1.0-397B60.8%estimated ± 3.6 pp, high confidence
43Gemini 3.7 Flash60.7%estimated ± 1.6 pp, medium confidence
44GPT-6.1 Sol60.4%estimated ± 5.3 pp, medium confidence
45Mistral Large 460.4%estimated ± 5.3 pp, medium confidence
46Ling 3.1 Flash60.2%estimated ± 5.3 pp, medium confidence
47Muse Spark 1.160.0%estimated ± 3.6 pp, high confidence
48Qwen3.8 Max Preview59.7%estimated ± 4.4 pp, high confidence
49SWE-1.759.4%estimated ± 1.6 pp, low confidence
50dots3-note Preview59.4%estimated ± 3.6 pp, high confidence
51GPT-5.1-Codex59.1%estimated ± 7.2 pp, low confidence
52GPT-5.1-Codex-Max59.1%estimated ± 7.2 pp, low confidence
53GLM-4.559.1%estimated ± 7.2 pp, low confidence
54GLM-4.659.1%estimated ± 7.2 pp, low confidence
55Grok Code Fast 159.1%estimated ± 7.2 pp, low confidence
56Qwen3.7 Max58.9%estimated ± 3.6 pp, high confidence
57GPT-5.558.4%measured
58Atria Dawn Preview57.6%estimated ± 3.6 pp, high confidence
59Ornith-1.5-35B-A3B57.6%estimated ± 3.6 pp, high confidence
60Laguna S 2.157.3%estimated ± 3.6 pp, high confidence
61MiniMax M356.8%estimated ± 3.6 pp, high confidence
62Sakana Fugu56.8%estimated ± 3.6 pp, high confidence
63Kimi K2.656.2%estimated ± 3.6 pp, high confidence
64Composer 2.556.1%measured
65GLM-5.155.9%estimated ± 3.6 pp, high confidence
66Gemini 3.1 Pro55.8%estimated ± 4.4 pp, high confidence
67Claude Opus 4.755.6%estimated ± 1.6 pp, low confidence
68Gemini 3 Flash55.5%estimated ± 5.4 pp, low confidence
69GLM-5.255.0%measured
70GPT-5.454.9%estimated ± 3.6 pp, high confidence
71Qwen3.7 Plus54.7%estimated ± 3.6 pp, high confidence
72Qwen 3.6 Max (preview)54.3%estimated ± 3.6 pp, high confidence
73MiMo-V2.5-Pro54.1%estimated ± 3.6 pp, high confidence
74GPT-5.2-Codex54.1%estimated ± 5.4 pp, low confidence
75Claude Opus 4.553.9%estimated ± 3.6 pp, high confidence
76Gemini 3.6 Flash53.5%measured
77GPT-5.3 Codex53.5%estimated ± 3.6 pp, high confidence
78Ling 3.0 Flash53.1%estimated ± 3.6 pp, high confidence
79Qwen3.6 Plus53.1%estimated ± 3.6 pp, high confidence
80Step 3.7 Flash52.6%estimated ± 3.6 pp, high confidence
81MiniMax M2.752.5%estimated ± 3.6 pp, high confidence
82MiMo-V2.552.3%estimated ± 3.6 pp, high confidence
83Inkling-Small52.0%estimated ± 3.6 pp, high confidence
84GPT-5.251.4%estimated ± 3.6 pp, high confidence
85DeepSeek V4 Pro 081351.1%estimated ± 3.6 pp, high confidence
86GLM-550.6%estimated ± 3.6 pp, high confidence
87Kimi K2.7 Code49.7%measured
88Qwen3.5 Flash49.6%estimated ± 5.4 pp, low confidence
89Inkling49.1%estimated ± 3.6 pp, medium confidence
90Gemini 3.5 Flash-Lite48.9%estimated ± 3.6 pp, medium confidence
91Gemini 3.5 Flash48.8%measured
92Gemini 3.1 Flash-Lite48.6%estimated ± 5.4 pp, low confidence
93Qwen3.6-27B47.6%estimated ± 3.6 pp, medium confidence
94Quasar 438B46.5%estimated ± 4.4 pp, high confidence
95MAI-Thinking-146.2%estimated ± 3.6 pp, medium confidence
96Apodex 1.146.0%estimated ± 4.4 pp, high confidence
97Apodex 1.1 Mini46.0%estimated ± 4.4 pp, high confidence
98Muse Spark45.4%estimated ± 3.6 pp, medium confidence
99Solar Pro 445.1%estimated ± 5.3 pp, low confidence
100Ling 3.0 Flash VL44.4%estimated ± 5.3 pp, low confidence
101Grok 4.2044.1%estimated ± 3.6 pp, medium confidence
102Hy343.9%estimated ± 4.4 pp, medium confidence
103Hy3 Preview43.9%estimated ± 4.4 pp, medium confidence
104Muse Glimmer 30B42.8%estimated ± 3.6 pp, medium confidence
105GPT-5.4 mini42.5%estimated ± 1.6 pp, low confidence
106Claude Opus 4.642.4%estimated ± 1.6 pp, low confidence
107Qwen3.5 397B42.2%estimated ± 3.6 pp, medium confidence
108Kimi K2.541.8%estimated ± 3.6 pp, medium confidence
109Ornith-1.0-35B41.1%estimated ± 3.6 pp, medium confidence
110GPT-5.4 nano41.0%estimated ± 4.4 pp, medium confidence
111K-EXAONE 2.040.9%estimated ± 5.3 pp, low confidence
112A.X K239.3%estimated ± 5.3 pp, low confidence
113Claude Sonnet 4.639.1%estimated ± 1.6 pp, low confidence
114Qwen3.6-35B-A3B39.1%estimated ± 3.6 pp, medium confidence
115Laguna M.138.4%estimated ± 3.6 pp, medium confidence
116DeepSeek V3 032436.1%estimated ± 5.3 pp, low confidence
117North Mini Code35.8%estimated ± 5.3 pp, low confidence
118Ling 3.0 Flash FP835.5%estimated ± 4.4 pp, medium confidence
119Mercury 2.535.3%estimated ± 5.3 pp, low confidence
120MiMo-V2-Flash34.7%estimated ± 4.4 pp, medium confidence
121Laguna XS 2.134.7%estimated ± 3.6 pp, medium confidence
122Ornith-1.5-9B34.5%estimated ± 3.6 pp, medium confidence
123GPT-5.134.3%estimated ± 4.4 pp, medium confidence
124Nemotron 3 Ultra34.2%estimated ± 4.4 pp, medium confidence
125Mistral Medium 3.5 128B32.0%estimated ± 4.4 pp, medium confidence
126Kimi K2.5 (Reasoning)31.9%estimated ± 4.4 pp, medium confidence
127Laguna XS.231.6%estimated ± 3.6 pp, medium confidence
128Qwen3.5-122B-A10B30.9%estimated ± 4.4 pp, medium confidence
129GLM-4.730.5%estimated ± 4.4 pp, medium confidence
130Gemma 4 31B28.9%estimated ± 4.4 pp, medium confidence
131Grok 4.327.9%estimated ± 4.4 pp, medium confidence
132MiMo-V2.6-Pro27.5%estimated ± 3.6 pp, low confidence
133MiMo-V2.6-Flash27.1%estimated ± 3.6 pp, low confidence
134o125.8%estimated ± 4.4 pp, medium confidence
135Gemma 4 26B A4B25.4%estimated ± 4.4 pp, medium confidence
136GPT-5 (high)24.2%estimated ± 4.4 pp, medium confidence
137Nemotron 3 Super 100B24.1%estimated ± 4.4 pp, medium confidence
138Ornith-1.0-9B23.7%estimated ± 3.6 pp, medium confidence
139DeepSeek V4.1 Flash21.8%estimated ± 3.6 pp, low confidence
140o1-preview21.3%estimated ± 4.4 pp, medium confidence
141Gemini 2.5 Pro20.7%estimated ± 4.4 pp, medium confidence
142K-Exaone19.8%estimated ± 4.4 pp, medium confidence
143Gemma 4 12B18.9%estimated ± 4.4 pp, medium confidence
144LongCat-Flash-Lite-Sparse18.7%estimated ± 3.6 pp, medium confidence
145GPT-OSS 120B18.6%estimated ± 4.4 pp, medium confidence
146Command A+16.7%estimated ± 4.4 pp, medium confidence
147Nemotron 3.5 Lightning 30B A3B NVFP415.9%estimated ± 4.4 pp, medium confidence
148Mistral Small 415.9%estimated ± 4.4 pp, medium confidence
149Mistral Small 4 (Reasoning)15.9%estimated ± 4.4 pp, medium confidence
150Trinity-Large-Preview15.3%estimated ± 4.4 pp, medium confidence
151Trinity-Large-Thinking15.3%estimated ± 4.4 pp, medium confidence
152Ling 2.6 Flash14.9%estimated ± 4.4 pp, medium confidence
153Solar Pro 314.6%estimated ± 5.3 pp, low confidence
154Granite 4.2 3B14.3%estimated ± 5.3 pp, low confidence
155Gemini 1.5 Pro13.8%estimated ± 4.4 pp, medium confidence
156DeepSeek V313.4%estimated ± 4.4 pp, medium confidence
157Ling 3.0 Tiny12.5%estimated ± 5.3 pp, low confidence
158GPT-4 Turbo12.4%estimated ± 4.4 pp, medium confidence
159GPT-OSS 20B11.9%estimated ± 4.4 pp, medium confidence
160GPT-4.1 mini11.6%estimated ± 4.4 pp, medium confidence
161Mistral Large 311.5%estimated ± 4.4 pp, medium confidence
162Claude 3 Opus11.1%estimated ± 4.4 pp, medium confidence
163Llama 4 Maverick9.1%estimated ± 4.4 pp, medium confidence
164Celeris-18.0%estimated ± 4.4 pp, medium confidence
165Nemotron 3 Nano 30B7.9%estimated ± 4.4 pp, medium confidence
166Nemotron 3 Nano Omni 30B A3B7.6%estimated ± 4.4 pp, medium confidence
167Granite 4.2 30B6.9%estimated ± 3.6 pp, medium confidence
168Ultravox v0.6 Llama 3.3 70B6.5%estimated ± 4.4 pp, medium confidence
169GPT-4o mini6.2%estimated ± 4.4 pp, medium confidence
170GPT-4.1 nano6.0%estimated ± 4.4 pp, medium confidence
171Gemma 3 27B5.4%estimated ± 4.4 pp, medium confidence
172Gemma 4 E4B5.0%estimated ± 4.4 pp, medium confidence
173Llama 4 Scout4.4%estimated ± 4.4 pp, medium confidence
174LFM2.5-2.6B4.1%estimated ± 4.4 pp, medium confidence
175LLaDA2.2-flash3.9%estimated ± 3.6 pp, medium confidence
176Gemma 4 E2B3.8%estimated ± 4.4 pp, medium confidence
177Granite 4.2 8B0.3%estimated ± 3.6 pp, medium confidence
178MiniCPM5-2B0.0%estimated ± 3.6 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General