benchgap
Coding

LiveCodeBench leaderboard

As of 2026-10-07, the highest measured score on LiveCodeBench is 91.6% by Qwen3.7 Max. 140 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Apodex 1.1100.0%estimated ± 8.9 pp, low confidence
2Apodex 1.1 Mini100.0%estimated ± 8.9 pp, low confidence
3Claude Fable 5100.0%estimated ± 8.9 pp, low confidence
4Claude Fable 5.1100.0%estimated ± 8.9 pp, low confidence
5Claude Opus 4.7 (Adaptive)100.0%estimated ± 8.9 pp, low confidence
6Claude Opus 4.8100.0%estimated ± 8.9 pp, low confidence
7Claude Opus 5100.0%estimated ± 8.9 pp, low confidence
8Claude Sonnet 5100.0%estimated ± 8.9 pp, low confidence
9DeepSeek V4 Flash 0731100.0%estimated ± 8.9 pp, low confidence
10DeepSeek V4 Pro 0813100.0%estimated ± 8.9 pp, low confidence
11Gemini 3.1 Pro100.0%estimated ± 8.9 pp, low confidence
12Gemini 3.5 Flash100.0%estimated ± 8.9 pp, low confidence
13Gemini 3.6 Flash100.0%estimated ± 8.9 pp, low confidence
14Gemini 3.7 Flash100.0%estimated ± 8.9 pp, low confidence
15Gemini 3.8 Flash100.0%estimated ± 8.9 pp, low confidence
16GLM-5.1100.0%estimated ± 8.9 pp, low confidence
17GLM-5.2100.0%estimated ± 8.9 pp, low confidence
18GLM-5.3100.0%estimated ± 8.9 pp, low confidence
19GPT-5.4100.0%estimated ± 8.9 pp, low confidence
20GPT-5.4 mini100.0%estimated ± 8.9 pp, low confidence
21GPT-5.4 nano100.0%estimated ± 8.9 pp, low confidence
22GPT-5.5100.0%estimated ± 8.9 pp, low confidence
23GPT-5.6 Luna100.0%estimated ± 8.9 pp, low confidence
24GPT-5.6 Sol100.0%estimated ± 8.9 pp, low confidence
25GPT-5.6 Terra100.0%estimated ± 8.9 pp, low confidence
26GPT-6 Astra100.0%estimated ± 8.9 pp, low confidence
27Grok 4.5100.0%estimated ± 8.9 pp, low confidence
28Grok 4.6100.0%estimated ± 8.9 pp, low confidence
29Hy3100.0%estimated ± 8.9 pp, low confidence
30Hy3 Preview100.0%estimated ± 8.9 pp, low confidence
31Inkling100.0%estimated ± 8.9 pp, low confidence
32Inkling-Small100.0%estimated ± 8.9 pp, low confidence
33Kimi K2.6100.0%estimated ± 8.9 pp, low confidence
34Kimi K2.7 Code100.0%estimated ± 8.9 pp, low confidence
35Kimi K3100.0%estimated ± 8.9 pp, low confidence
36Ling 3.0 Flash100.0%estimated ± 8.9 pp, low confidence
37Ling 3.0 Flash FP8100.0%estimated ± 8.9 pp, low confidence
38MiMo-V2.5-Pro100.0%estimated ± 8.9 pp, low confidence
39MiniMax M2.7100.0%estimated ± 8.9 pp, low confidence
40MiniMax M3100.0%estimated ± 8.9 pp, low confidence
41Muse Spark100.0%estimated ± 8.9 pp, low confidence
42Muse Spark 1.1100.0%estimated ± 8.9 pp, low confidence
43Muse Spark 1.2100.0%estimated ± 8.9 pp, low confidence
44Muse Spark 1.3100.0%estimated ± 8.9 pp, low confidence
45Quasar 438B100.0%estimated ± 8.9 pp, low confidence
46Qwen3.6 Plus100.0%estimated ± 8.9 pp, low confidence
47Qwen3.8-27B100.0%estimated ± 8.9 pp, low confidence
48Qwen3.8-Flash-Next100.0%estimated ± 8.9 pp, low confidence
49Qwen3.8 Max Preview100.0%estimated ± 8.9 pp, low confidence
50MiMo-V2-Flash98.4%estimated ± 8.9 pp, low confidence
51GPT-5.197.2%estimated ± 8.9 pp, low confidence
52Gemini 3.5 Flash-Lite97.1%estimated ± 8.9 pp, low confidence
53Nemotron 3 Ultra96.9%estimated ± 8.9 pp, low confidence
54Muse Glimmer 30B96.2%estimated ± 8.9 pp, low confidence
55Claude Mythos 595.3%estimated ± 10.2 pp, low confidence
56Ember-193.5%estimated ± 10.2 pp, low confidence
57Qwen3.7 Max91.6%measured
58Mistral Medium 3.5 128B90.8%estimated ± 8.9 pp, low confidence
59Kimi K2.590.5%estimated ± 8.9 pp, low confidence
60Kimi K2.5 (Reasoning)90.5%estimated ± 8.9 pp, low confidence
61Ornith-1.5-397B90.1%estimated ± 10.2 pp, low confidence
62Qwen3.7 Plus89.6%measured
63GPT-5.3 Codex89.5%estimated ± 10.2 pp, low confidence
64Ornith-1.0-397B87.9%estimated ± 10.2 pp, low confidence
65Solar Pro 487.8%measured
66Qwen3.5-122B-A10B87.7%estimated ± 8.9 pp, low confidence
67Claude Opus 4.587.0%estimated ± 10.2 pp, low confidence
68Beam87.0%estimated ± 10.2 pp, low confidence
69Claude Opus 4.687.0%estimated ± 10.2 pp, low confidence
70GPT-5.286.5%estimated ± 10.2 pp, low confidence
71Claude Sonnet 4.686.2%estimated ± 10.2 pp, low confidence
72Ornith-1.5-35B-A3B85.9%estimated ± 10.2 pp, low confidence
73BTL-485.5%estimated ± 10.2 pp, low confidence
74dots3-note Preview85.5%estimated ± 10.2 pp, low confidence
75MiMo-V2-Pro85.2%estimated ± 10.2 pp, low confidence
76GLM-585.1%estimated ± 10.2 pp, low confidence
77GLM-4.784.9%measured
78Claude Sonnet 4.584.7%estimated ± 10.2 pp, low confidence
79Grok 4.2084.4%estimated ± 10.2 pp, low confidence
80Qwen3.5 397B84.1%estimated ± 10.2 pp, low confidence
81Qwen3.6-27B83.9%measured
82Ornith-1.0-35B83.7%estimated ± 10.2 pp, low confidence
83MiMo-V2-Omni83.2%estimated ± 10.2 pp, low confidence
84Laguna M.183.1%estimated ± 10.2 pp, low confidence
85Claude 4.1 Opus83.0%estimated ± 10.2 pp, low confidence
86MAI-Thinking-182.4%estimated ± 10.2 pp, low confidence
87Claude Haiku 4.582.2%estimated ± 10.2 pp, low confidence
88Gemma 4 31B82.1%estimated ± 8.9 pp, low confidence
89Claude 4 Sonnet81.8%estimated ± 10.2 pp, low confidence
90MAI-Code-1.1-Flash81.8%estimated ± 10.2 pp, low confidence
91Qwen3.5-27B81.6%estimated ± 10.2 pp, low confidence
92Laguna XS 2.180.6%estimated ± 10.2 pp, low confidence
93Grok Code Fast 180.5%estimated ± 10.2 pp, low confidence
94Ornith-1.5-9B80.4%estimated ± 10.2 pp, low confidence
95Qwen3.6-35B-A3B80.4%measured
96Solar Open 280.3%estimated ± 10.2 pp, low confidence
97Laguna XS.279.9%estimated ± 10.2 pp, low confidence
98Ornith-1.0-9B79.6%estimated ± 10.2 pp, low confidence
99Qwen3.5-35B-A3B79.4%estimated ± 10.2 pp, low confidence
100Grok 4.379.2%estimated ± 8.9 pp, low confidence
101K-EXAONE 2.078.8%estimated ± 10.2 pp, low confidence
102LongCat-Flash-Lite-Sparse78.8%estimated ± 10.2 pp, low confidence
103Ternary Bonsai 2 27B73.3%estimated ± 10.2 pp, low confidence
104o173.2%estimated ± 8.9 pp, low confidence
105Step 3.7 Flash72.9%estimated ± 8.9 pp, low confidence
106Gemma 4 26B A4B72.3%estimated ± 8.9 pp, low confidence
107Granite 4.2 30B70.4%estimated ± 10.2 pp, low confidence
108GPT-5 (high)68.7%estimated ± 8.9 pp, low confidence
109Nemotron 3 Super 100B68.6%estimated ± 8.9 pp, low confidence
110GPT-4.168.4%estimated ± 10.2 pp, low confidence
111ZAYA1-74B-Preview67.3%estimated ± 10.2 pp, low confidence
112o3-mini63.9%estimated ± 10.2 pp, low confidence
113LLaDA2.2-flash63.9%estimated ± 10.2 pp, low confidence
114Claude 3.5 Sonnet63.6%estimated ± 10.2 pp, low confidence
115MiniCPM5-2B61.3%estimated ± 10.2 pp, low confidence
116o1-preview60.4%estimated ± 8.9 pp, low confidence
117Gemini 2.5 Pro58.7%estimated ± 8.9 pp, low confidence
118K-Exaone56.3%estimated ± 8.9 pp, low confidence
119Gemma 4 12B53.8%estimated ± 8.9 pp, low confidence
120GPT-OSS 120B52.7%estimated ± 8.9 pp, low confidence
121Command A+47.4%estimated ± 8.9 pp, low confidence
122Nemotron 3.5 Lightning 30B A3B NVFP445.2%estimated ± 8.9 pp, low confidence
123Mistral Small 445.0%estimated ± 8.9 pp, low confidence
124Mistral Small 4 (Reasoning)45.0%estimated ± 8.9 pp, low confidence
125Trinity-Large-Preview43.3%estimated ± 8.9 pp, low confidence
126Trinity-Large-Thinking43.3%estimated ± 8.9 pp, low confidence
127Ling 2.6 Flash42.2%estimated ± 8.9 pp, low confidence
128Gemini 1.5 Pro39.1%estimated ± 8.9 pp, low confidence
129DeepSeek V337.6%measured
130Granite 4.2 8B36.7%estimated ± 8.9 pp, low confidence
131GPT-4 Turbo35.0%estimated ± 8.9 pp, low confidence
132GPT-OSS 20B33.5%estimated ± 8.9 pp, low confidence
133GPT-4.1 mini32.6%estimated ± 8.9 pp, low confidence
134Mistral Large 332.3%estimated ± 8.9 pp, low confidence
135Claude 3 Opus31.3%estimated ± 8.9 pp, low confidence
136Llama 4 Maverick25.5%estimated ± 8.9 pp, low confidence
137Celeris-122.1%estimated ± 8.9 pp, low confidence
138Nemotron 3 Nano 30B22.1%estimated ± 8.9 pp, low confidence
139Nemotron 3 Nano Omni 30B A3B21.0%estimated ± 8.9 pp, low confidence
140Ultravox v0.6 Llama 3.3 70B17.9%estimated ± 8.9 pp, low confidence
141GPT-4o mini17.0%estimated ± 8.9 pp, low confidence
142GPT-4.1 nano16.6%estimated ± 8.9 pp, low confidence
143Gemma 3 27B14.7%estimated ± 8.9 pp, low confidence
144Gemma 4 E4B13.6%estimated ± 8.9 pp, low confidence
145Llama 4 Scout11.6%estimated ± 8.9 pp, low confidence
146LFM2.5-2.6B10.9%estimated ± 8.9 pp, low confidence
147Gemma 4 E2B10.1%estimated ± 8.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General