benchgap
Coding

LiveCodeBench v6 leaderboard

As of 2026-10-07, the highest measured score on LiveCodeBench v6 is 93.2% by Sakana Fugu-Ultra. 157 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.5100.0%estimated ± 2.1 pp, medium confidence
2Gemini 4 Argon100.0%estimated ± 3.3 pp, medium confidence
3MiMo-V2.6-Pro100.0%estimated ± 3.3 pp, medium confidence
4Step 5 Preview99.4%estimated ± 3.3 pp, medium confidence
5Claude Sonnet 5.599.0%estimated ± 2.1 pp, medium confidence
6Claude Fable 5.198.9%estimated ± 2.1 pp, medium confidence
7GPT-6 Sol98.2%estimated ± 3.3 pp, medium confidence
8Grok 4.798.0%estimated ± 3.3 pp, medium confidence
9Claude Mythos 596.2%estimated ± 2.1 pp, high confidence
10Claude Fable 596.1%estimated ± 2.1 pp, high confidence
11Claude Haiku 5.595.9%estimated ± 3.3 pp, medium confidence
12Claude Opus 595.8%estimated ± 2.1 pp, high confidence
13GPT-6 Luna95.5%estimated ± 3.3 pp, medium confidence
14GPT-6.1 Sol95.2%estimated ± 3.3 pp, medium confidence
15Mistral Large 495.2%estimated ± 3.3 pp, medium confidence
16Ling 3.1 Flash95.1%estimated ± 3.3 pp, medium confidence
17GPT-6 Astra93.6%estimated ± 3.1 pp, medium confidence
18Grok 4.693.6%estimated ± 3.1 pp, medium confidence
19Gemini 3.8 Flash93.5%estimated ± 3.1 pp, medium confidence
20Kimi K393.5%estimated ± 3.1 pp, medium confidence
21Gemini 3.7 Flash93.4%estimated ± 3.1 pp, medium confidence
22Muse Spark 1.393.4%estimated ± 3.1 pp, medium confidence
23Sakana Fugu-Ultra93.2%measured
24GLM-5.393.1%estimated ± 3.1 pp, medium confidence
25DeepSeek V4.1 Flash93.1%estimated ± 3.3 pp, medium confidence
26Qwen3.8 Max93.0%estimated ± 2.1 pp, high confidence
27Sakana Fugu92.9%measured
28GLM-5.3-Flash92.8%estimated ± 3.3 pp, medium confidence
29Qwen3.8-Omni-Flash92.6%measured
30MiMo-V2.6-Flash92.5%estimated ± 3.3 pp, high confidence
31Muse Spark 1.292.5%estimated ± 3.1 pp, high confidence
32Solar Open 292.4%measured
33Qwen3.8 Max Preview92.3%estimated ± 3.1 pp, high confidence
34Claude Opus 4.892.2%estimated ± 2.1 pp, high confidence
35Hy4 preview92.1%estimated ± 2.1 pp, high confidence
36Qwen3.8-Flash-Next91.9%measured
37Gemini 3.6 Flash91.7%estimated ± 3.1 pp, high confidence
38Grok 4.591.6%estimated ± 2.1 pp, high confidence
39GPT-5.6 Sol91.6%estimated ± 2.1 pp, high confidence
40Gemini 3.1 Pro91.5%estimated ± 3.1 pp, high confidence
41dots3-note Preview91.5%measured
42GPT-5.6 Terra91.1%estimated ± 2.1 pp, high confidence
43Beam90.9%estimated ± 2.1 pp, high confidence
44GPT-5.6 Luna90.8%estimated ± 2.1 pp, high confidence
45Ornith-1.5-397B90.7%estimated ± 2.1 pp, high confidence
46Seed 2.1 Pro90.6%estimated ± 3.8 pp, low confidence
47GLM-5.290.5%estimated ± 2.1 pp, high confidence
48Claude Opus 4.7 (Adaptive)90.4%estimated ± 2.1 pp, high confidence
49Qwen3.8-27B90.3%measured
50Muse Spark 1.190.2%estimated ± 2.1 pp, high confidence
51Ternary Bonsai 2 27B90.1%measured
52Claude Sonnet 590.0%estimated ± 2.1 pp, high confidence
53Ornith-1.0-397B89.6%estimated ± 2.1 pp, high confidence
54Kimi K2.689.6%measured
55Atria Dawn Preview89.4%estimated ± 2.1 pp, high confidence
56Quasar 438B89.3%estimated ± 3.1 pp, high confidence
57Laguna S 2.189.3%estimated ± 2.1 pp, high confidence
58Apodex 1.189.2%estimated ± 3.1 pp, high confidence
59Apodex 1.1 Mini89.2%estimated ± 3.1 pp, high confidence
60Kimi K2.7 Code89.2%estimated ± 3.1 pp, high confidence
61Qwen3.7 Max89.0%estimated ± 2.1 pp, high confidence
62Nemotron 3 Ultra89.0%measured
63GPT-5.589.0%estimated ± 2.1 pp, high confidence
64GLM-5.188.9%estimated ± 2.1 pp, high confidence
65Ornith-1.5-35B-A3B88.6%estimated ± 2.1 pp, high confidence
66Hy388.6%estimated ± 3.1 pp, high confidence
67Hy3 Preview88.6%estimated ± 3.1 pp, high confidence
68GPT-5.488.6%estimated ± 2.1 pp, high confidence
69Seed 2.1 Turbo88.5%estimated ± 3.8 pp, low confidence
70MiniMax M388.4%estimated ± 2.1 pp, high confidence
71Qwen 3.6 Max (preview)88.4%estimated ± 2.1 pp, high confidence
72MiMo-V2.5-Pro88.3%estimated ± 2.1 pp, high confidence
73BTL-388.1%measured
74Ling 3.0 Flash88.1%estimated ± 2.1 pp, high confidence
75Step 3.7 Flash87.9%estimated ± 2.1 pp, high confidence
76MiniMax M2.787.9%estimated ± 2.1 pp, high confidence
77MiMo-V2.587.9%estimated ± 2.1 pp, high confidence
78Qwen3.7 Plus87.8%estimated ± 2.1 pp, high confidence
79MAI-Thinking-187.7%measured
80GPT-5.4 mini87.7%estimated ± 3.1 pp, high confidence
81GPT-5.4 nano87.7%estimated ± 3.1 pp, high confidence
82GPT-5.3 Codex87.5%estimated ± 2.1 pp, high confidence
83Gemini 3.5 Flash87.4%estimated ± 2.1 pp, high confidence
84Inkling-Small87.2%estimated ± 2.1 pp, high confidence
85Ember-187.1%estimated ± 8.1 pp, low confidence
86Qwen3.6 Plus87.1%measured
87GPT-5.287.1%estimated ± 2.1 pp, high confidence
88Gemini 3.5 Flash-Lite87.0%estimated ± 2.1 pp, high confidence
89DeepSeek V4 Pro 081387.0%estimated ± 2.1 pp, high confidence
90GLM-586.9%estimated ± 2.1 pp, high confidence
91Inkling86.6%estimated ± 2.1 pp, high confidence
92Solar Pro 486.5%estimated ± 3.3 pp, high confidence
93Qwen3.6-27B86.2%estimated ± 2.1 pp, high confidence
94Claude Opus 4.686.2%estimated ± 2.1 pp, high confidence
95Claude Sonnet 4.686.2%estimated ± 8.1 pp, low confidence
96Ling 3.0 Flash VL86.1%estimated ± 3.3 pp, high confidence
97MiMo-V2-Pro86.0%estimated ± 8.1 pp, low confidence
98Claude Sonnet 4.585.9%estimated ± 8.1 pp, low confidence
99DeepSeek V4 Flash 073185.9%estimated ± 2.1 pp, high confidence
100Muse Spark85.8%estimated ± 2.1 pp, high confidence
101Ling 3.0 Flash FP885.7%estimated ± 3.1 pp, high confidence
102Grok 4.2085.6%estimated ± 2.1 pp, high confidence
103MiMo-V2-Omni85.5%estimated ± 8.1 pp, low confidence
104Claude 4.1 Opus85.4%estimated ± 8.1 pp, low confidence
105MiMo-V2-Flash85.4%estimated ± 3.1 pp, high confidence
106Muse Glimmer 30B85.3%estimated ± 2.1 pp, high confidence
107GPT-5.185.2%estimated ± 3.1 pp, high confidence
108Claude Haiku 4.585.2%estimated ± 8.1 pp, low confidence
109Claude 4 Sonnet85.1%estimated ± 8.1 pp, low confidence
110MAI-Code-1.1-Flash85.1%estimated ± 8.1 pp, low confidence
111Qwen3.5-27B85.0%estimated ± 8.1 pp, low confidence
112Ornith-1.0-35B85.0%estimated ± 2.1 pp, high confidence
113Kimi K2.585.0%measured
114Claude Opus 4.584.8%measured
115Grok Code Fast 184.7%estimated ± 8.1 pp, low confidence
116Qwen3.6-35B-A3B84.6%estimated ± 2.1 pp, high confidence
117Laguna M.184.5%estimated ± 2.1 pp, high confidence
118Mistral Medium 3.5 128B84.3%estimated ± 3.1 pp, high confidence
119Qwen3.5-35B-A3B84.3%estimated ± 8.1 pp, low confidence
120Kimi K2.5 (Reasoning)84.2%estimated ± 3.1 pp, high confidence
121K-EXAONE 2.084.1%estimated ± 3.3 pp, high confidence
122A.X K284.0%measured
123Laguna XS 2.183.9%estimated ± 2.1 pp, high confidence
124Ornith-1.5-9B83.8%estimated ± 2.1 pp, high confidence
125Qwen3.5-122B-A10B83.8%estimated ± 3.1 pp, high confidence
126GLM-4.783.6%estimated ± 3.1 pp, high confidence
127Qwen3.5 397B83.6%measured
128Laguna XS.283.3%estimated ± 2.1 pp, high confidence
129Gemma 4 31B82.8%estimated ± 3.1 pp, high confidence
130Grok 4.382.3%estimated ± 3.1 pp, high confidence
131Ornith-1.0-9B81.9%estimated ± 2.1 pp, high confidence
132DeepSeek V3 032481.4%estimated ± 3.3 pp, high confidence
133North Mini Code81.3%estimated ± 3.3 pp, high confidence
134o181.2%estimated ± 3.1 pp, high confidence
135Gemma 4 26B A4B81.0%estimated ± 3.1 pp, high confidence
136LongCat-Flash-Lite-Sparse81.0%estimated ± 2.1 pp, high confidence
137Mercury 2.581.0%estimated ± 3.3 pp, high confidence
138GPT-5 (high)80.3%estimated ± 3.1 pp, high confidence
139Nemotron 3 Super 100B80.2%estimated ± 3.1 pp, high confidence
140o1-preview78.4%estimated ± 3.1 pp, high confidence
141Gemini 2.5 Pro78.0%estimated ± 3.1 pp, high confidence
142K-Exaone77.4%estimated ± 3.1 pp, high confidence
143GPT-4.176.7%estimated ± 8.1 pp, low confidence
144LLaDA2.2-flash76.5%estimated ± 2.1 pp, high confidence
145GPT-OSS 120B76.5%estimated ± 3.1 pp, high confidence
146Granite 4.2 30B75.8%measured
147Command A+75.0%estimated ± 3.1 pp, high confidence
148Nemotron 3.5 Lightning 30B A3B NVFP474.3%estimated ± 3.1 pp, high confidence
149Mistral Small 474.2%estimated ± 3.1 pp, high confidence
150Mistral Small 4 (Reasoning)74.2%estimated ± 3.1 pp, high confidence
151Trinity-Large-Preview73.7%estimated ± 3.1 pp, high confidence
152Trinity-Large-Thinking73.7%estimated ± 3.1 pp, high confidence
153Ling 2.6 Flash73.4%estimated ± 3.1 pp, high confidence
154Granite 4.2 8B73.2%measured
155Gemini 1.5 Pro72.3%estimated ± 3.1 pp, high confidence
156Gemma 4 12B72.0%measured
157DeepSeek V371.9%estimated ± 3.1 pp, high confidence
158o3-mini71.7%estimated ± 8.1 pp, low confidence
159Claude 3.5 Sonnet71.4%estimated ± 8.1 pp, low confidence
160GPT-4 Turbo70.8%estimated ± 3.1 pp, high confidence
161GPT-OSS 20B70.3%estimated ± 3.1 pp, high confidence
162GPT-4.1 mini69.9%estimated ± 3.1 pp, high confidence
163Mellum2-12B-A2.5B-Thinking69.9%measured
164Mistral Large 369.8%estimated ± 3.1 pp, high confidence
165Granite 4.2 3B69.7%measured
166Claude 3 Opus69.5%estimated ± 3.1 pp, high confidence
167Solar Pro 369.2%estimated ± 3.3 pp, high confidence
168MiniCPM5-2B69.1%measured
169Ling 3.0 Tiny68.1%estimated ± 3.3 pp, high confidence
170Llama 4 Maverick67.0%estimated ± 3.1 pp, high confidence
171BTL-466.1%measured
172ZAYA1-8B65.8%measured
173ZAYA1-74B-Preview65.7%measured
174Celeris-165.4%estimated ± 3.1 pp, high confidence
175Nemotron 3 Nano 30B65.4%estimated ± 3.1 pp, high confidence
176Nemotron 3 Nano Omni 30B A3B64.9%estimated ± 3.1 pp, high confidence
177Ultravox v0.6 Llama 3.3 70B63.3%estimated ± 3.1 pp, high confidence
178GPT-4o mini62.8%estimated ± 3.1 pp, high confidence
179GPT-4.1 nano62.6%estimated ± 3.1 pp, high confidence
180Gemma 3 27B61.6%estimated ± 3.1 pp, high confidence
181Gemma 4 E4B60.9%estimated ± 3.1 pp, high confidence
182Llama 4 Scout59.7%estimated ± 3.1 pp, high confidence
183Agents-A1-4B59.6%measured
184LFM2.5-2.6B59.4%measured
185Gemma 4 E2B58.8%estimated ± 3.1 pp, medium confidence
186Mellum2-12B-A2.5B-Instruct37.2%measured
187MiniCPM5-1B33.5%measured
188LLaDA2.2-mini28.1%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General