benchgap
Coding

Vibe Code Bench leaderboard

As of 2026-10-07, the highest measured score on Vibe Code Bench is 91.9% by Gemini 4 Argon. 135 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5.1100.0%estimated ± 10.3 pp, low confidence
2Gemini 4 Argon91.9%measured
3Claude Opus 5.581.7%estimated ± 5.6 pp, low confidence
4Claude Fable 581.0%estimated ± 5.6 pp, low confidence
5Claude Opus 580.9%estimated ± 5.6 pp, low confidence
6GPT-6 Astra80.8%estimated ± 5.6 pp, low confidence
7Mistral Large 478.4%measured
8SWE-278.2%estimated ± 5.6 pp, low confidence
9GPT-5.6 Sol77.7%estimated ± 10.3 pp, low confidence
10Claude Opus 4.875.2%estimated ± 5.6 pp, low confidence
11Claude Haiku 5.575.2%estimated ± 5.6 pp, low confidence
12Grok 4.675.0%estimated ± 10.3 pp, low confidence
13Claude Sonnet 5.575.0%estimated ± 5.6 pp, low confidence
14GPT-5.6 Terra74.5%estimated ± 10.3 pp, low confidence
15Gemini 3.8 Flash72.9%estimated ± 10.3 pp, low confidence
16Kimi K372.7%estimated ± 10.3 pp, low confidence
17Gemini 3.7 Flash72.6%estimated ± 5.6 pp, low confidence
18Claude Sonnet 571.8%estimated ± 5.6 pp, low confidence
19SWE-1.771.4%estimated ± 5.6 pp, low confidence
20Claude Opus 4.771.0%measured
21Muse Spark 1.370.9%estimated ± 10.3 pp, low confidence
22GPT-5.569.8%measured
23GPT-5.467.4%measured
24GLM-5.367.0%estimated ± 10.3 pp, low confidence
25Composer 266.5%estimated ± 6.5 pp, low confidence
26Composer 2 Fast66.5%estimated ± 6.5 pp, low confidence
27Claude Mythos 562.5%estimated ± 8.1 pp, low confidence
28Ember-162.5%estimated ± 8.1 pp, low confidence
29Claude Opus 4.7 (Adaptive)62.4%estimated ± 8.1 pp, low confidence
30Ornith-1.5-397B62.3%estimated ± 8.1 pp, low confidence
31GPT-5.3 Codex61.8%measured
32Qwen3.8-Flash-Next61.2%estimated ± 10.3 pp, low confidence
33Ornith-1.0-397B59.8%estimated ± 8.1 pp, medium confidence
34Grok 4.559.4%estimated ± 10.3 pp, low confidence
35Muse Spark 1.258.7%estimated ± 10.3 pp, low confidence
36Claude Opus 4.657.6%measured
37Qwen3.8 Max Preview57.5%estimated ± 10.3 pp, low confidence
38GPT-5.6 Luna56.5%estimated ± 10.3 pp, low confidence
39Muse Spark 1.156.1%estimated ± 10.3 pp, low confidence
40Claude Opus 4.554.1%estimated ± 8.1 pp, medium confidence
41Beam54.1%estimated ± 8.1 pp, medium confidence
42GPT-5.253.5%measured
43Claude Opus 4.6 (Adaptive)53.5%measured
44Claude Sonnet 4.651.5%measured
45MiniMax M351.4%estimated ± 8.1 pp, medium confidence
46Gemini 3.6 Flash50.7%estimated ± 10.3 pp, low confidence
47Qwen3.7 Max50.6%estimated ± 8.1 pp, medium confidence
48DeepSeek V4 Pro 081349.9%measured
49GLM-5.249.5%estimated ± 10.3 pp, low confidence
50Inkling-Small48.9%estimated ± 8.1 pp, medium confidence
51Gemini 3.5 Flash48.7%measured
52Qwen3.8-27B48.0%estimated ± 10.3 pp, low confidence
53GPT-5.4 mini48.0%measured
54GPT-5.2-Codex37.9%measured
55Kimi K2.637.9%measured
56DeepSeek V4 Flash 073135.7%estimated ± 8.1 pp, medium confidence
57Ornith-1.5-35B-A3B35.7%estimated ± 8.1 pp, medium confidence
58Quasar 438B35.4%estimated ± 10.3 pp, low confidence
59Apodex 1.1 Mini34.7%estimated ± 10.3 pp, low confidence
60Kimi K2.7 Code34.7%estimated ± 10.3 pp, low confidence
61MiMo-V2.5-Pro33.9%estimated ± 10.3 pp, low confidence
62Gemini 3.1 Pro32.0%measured
63Hy332.0%estimated ± 10.3 pp, low confidence
64GLM-5.131.5%measured
65BTL-428.1%estimated ± 8.1 pp, medium confidence
66dots3-note Preview28.1%estimated ± 8.1 pp, medium confidence
67Kimi K2.527.9%estimated ± 6.5 pp, low confidence
68Gemma 4 31B27.7%estimated ± 6.5 pp, low confidence
69GLM-527.7%estimated ± 6.5 pp, low confidence
70Grok 427.7%estimated ± 6.5 pp, low confidence
71DeepSeek V3.227.7%estimated ± 6.5 pp, low confidence
72GPT-OSS 120B27.7%estimated ± 6.5 pp, low confidence
73GPT-OSS 20B27.7%estimated ± 6.5 pp, low confidence
74MiniMax M2.727.0%measured
75GPT-5.4 nano26.1%measured
76Qwen3.6 Plus25.6%measured
77GPT-5.124.6%measured
78GLM-5 (Reasoning)23.4%measured
79MiMo-V2-Pro23.1%estimated ± 8.1 pp, medium confidence
80Ling 3.0 Flash22.9%estimated ± 10.3 pp, low confidence
81Ling 3.0 Flash FP822.9%estimated ± 10.3 pp, low confidence
82Claude Sonnet 4.5 Thinking22.6%measured
83GPT-5.1-Codex-Max22.2%measured
84Gemini 3.5 Flash-Lite21.7%estimated ± 10.3 pp, low confidence
85Claude Opus 4.5 Thinking20.6%measured
86Gemini 3 Flash20.2%measured
87GPT-5 (high)20.1%measured
88Apodex 1.119.7%estimated ± 8.1 pp, medium confidence
89Qwen3.7 Plus19.7%estimated ± 8.1 pp, medium confidence
90Muse Spark19.7%measured
91Inkling18.6%estimated ± 8.1 pp, medium confidence
92Mistral Medium 3.5 128B18.6%estimated ± 8.1 pp, medium confidence
93Kimi K2.5 (Reasoning)17.5%measured
94Grok 4.316.3%estimated ± 10.3 pp, low confidence
95Qwen3.5 Plus15.7%measured
96MiniMax M2.514.9%measured
97Claude Sonnet 4.514.7%estimated ± 8.1 pp, medium confidence
98Qwen3.6-27B14.7%estimated ± 8.1 pp, medium confidence
99o114.7%estimated ± 10.3 pp, low confidence
100Step 3.7 Flash14.6%estimated ± 10.3 pp, low confidence
101Gemma 4 26B A4B14.4%estimated ± 10.3 pp, low confidence
102Gemini 3 Pro14.3%measured
103GPT-5 mini14.2%measured
104Nemotron 3 Super 100B13.5%estimated ± 10.3 pp, low confidence
105GPT-5.1-Codex13.1%measured
106o1-preview11.5%estimated ± 10.3 pp, low confidence
107Claude Haiku 4.5 Thinking11.4%measured
108K-Exaone10.5%estimated ± 10.3 pp, low confidence
109Gemma 4 12B10.0%estimated ± 10.3 pp, low confidence
110Command A+8.6%estimated ± 10.3 pp, low confidence
111Mistral Small 48.1%estimated ± 10.3 pp, low confidence
112Mistral Small 4 (Reasoning)8.1%estimated ± 10.3 pp, low confidence
113Trinity-Large-Preview7.7%estimated ± 10.3 pp, low confidence
114Trinity-Large-Thinking7.7%estimated ± 10.3 pp, low confidence
115Qwen3.5 397B7.7%estimated ± 8.1 pp, medium confidence
116Ling 2.6 Flash7.5%estimated ± 10.3 pp, low confidence
117Gemini 1.5 Pro6.9%estimated ± 10.3 pp, low confidence
118Muse Glimmer 30B6.7%estimated ± 8.1 pp, medium confidence
119GPT-4 Turbo6.1%estimated ± 10.3 pp, low confidence
120Mistral Large 35.6%estimated ± 10.3 pp, low confidence
121Claude 3 Opus5.4%estimated ± 10.3 pp, low confidence
122Ornith-1.0-35B5.1%estimated ± 8.1 pp, medium confidence
123DeepSeek V3.2 (Thinking)5.1%measured
124Llama 4 Maverick4.3%estimated ± 10.3 pp, low confidence
125Grok 4.204.1%measured
126Celeris-13.7%estimated ± 10.3 pp, low confidence
127Nemotron 3 Nano 30B3.7%estimated ± 10.3 pp, low confidence
128Nemotron 3 Nano Omni 30B A3B3.5%estimated ± 10.3 pp, low confidence
129Qwen3 Max3.5%measured
130GLM-4.63.1%measured
131MiMo-V2-Omni3.0%estimated ± 8.1 pp, medium confidence
132Ultravox v0.6 Llama 3.3 70B3.0%estimated ± 10.3 pp, low confidence
133GPT-4o mini2.9%estimated ± 10.3 pp, low confidence
134GPT-4.1 nano2.8%estimated ± 10.3 pp, low confidence
135Laguna M.12.7%estimated ± 8.1 pp, medium confidence
136Claude 4.1 Opus2.5%estimated ± 8.1 pp, medium confidence
137Gemma 3 27B2.5%estimated ± 10.3 pp, low confidence
138Hy3 Preview2.4%estimated ± 8.1 pp, medium confidence
139Gemma 4 E4B2.3%estimated ± 10.3 pp, low confidence
140Llama 4 Scout2.0%estimated ± 10.3 pp, low confidence
141LFM2.5-2.6B1.9%estimated ± 10.3 pp, low confidence
142Gemma 4 E2B1.7%estimated ± 10.3 pp, low confidence
143GLM-4.71.7%estimated ± 8.1 pp, medium confidence
144MAI-Thinking-11.5%estimated ± 8.1 pp, medium confidence
145MiMo-V2-Flash1.4%estimated ± 8.1 pp, medium confidence
146Qwen3.6-35B-A3B1.4%estimated ± 8.1 pp, medium confidence
147Claude Haiku 4.51.4%estimated ± 8.1 pp, medium confidence
148Grok 4.1 Fast (Reasoning)1.2%measured
149Claude 4 Sonnet1.1%estimated ± 8.1 pp, medium confidence
150MAI-Code-1.1-Flash1.0%estimated ± 8.1 pp, medium confidence
151Qwen3.5-27B1.0%estimated ± 8.1 pp, medium confidence
152Qwen3.5-122B-A10B0.9%estimated ± 8.1 pp, medium confidence
153Nemotron 3 Ultra0.9%estimated ± 8.1 pp, medium confidence
154Laguna XS 2.10.7%estimated ± 8.1 pp, medium confidence
155Grok Code Fast 10.7%estimated ± 8.1 pp, medium confidence
156Ornith-1.5-9B0.7%estimated ± 8.1 pp, medium confidence
157Solar Pro 40.7%estimated ± 8.1 pp, medium confidence
158Solar Open 20.7%estimated ± 8.1 pp, medium confidence
159Laguna XS.20.7%estimated ± 8.1 pp, medium confidence
160Ornith-1.0-9B0.7%estimated ± 8.1 pp, medium confidence
161Qwen3.5-35B-A3B0.7%estimated ± 8.1 pp, medium confidence
162K-EXAONE 2.00.7%estimated ± 8.1 pp, medium confidence
163LongCat-Flash-Lite-Sparse0.7%estimated ± 8.1 pp, medium confidence
164Claude 3.5 Sonnet0.6%estimated ± 8.1 pp, low confidence
165DeepSeek V30.6%estimated ± 8.1 pp, low confidence
166GPT-4.10.6%estimated ± 8.1 pp, low confidence
167GPT-4.1 mini0.6%estimated ± 8.1 pp, low confidence
168Granite 4.2 30B0.6%estimated ± 8.1 pp, low confidence
169Granite 4.2 8B0.6%estimated ± 8.1 pp, low confidence
170LLaDA2.2-flash0.6%estimated ± 8.1 pp, low confidence
171MiniCPM5-2B0.6%estimated ± 8.1 pp, low confidence
172Nemotron 3.5 Lightning 30B A3B NVFP40.6%estimated ± 8.1 pp, low confidence
173o3-mini0.6%estimated ± 8.1 pp, low confidence
174Ternary Bonsai 2 27B0.6%estimated ± 8.1 pp, low confidence
175ZAYA1-74B-Preview0.6%estimated ± 8.1 pp, low confidence
176Gemini 2.5 Pro0.4%measured
177Gemini 3.1 Flash-Lite0.0%measured
178Grok 4 Fast (Reasoning)0.0%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General