benchgap
Coding

Vals LiveCodeBench leaderboard

As of 2026-10-07, the highest measured score on Vals LiveCodeBench is 90.5% by Claude Fable 5.1. 94 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Ember-191.8%estimated ± 9.9 pp, low confidence
2Claude Fable 5.190.5%measured
3Claude Fable 589.8%measured
4Gemini 3.8 Flash89.5%measured
5Composer 289.4%estimated ± 3.8 pp, low confidence
6Claude Opus 5.589.3%estimated ± 1.9 pp, medium confidence
7Composer 2 Fast89.1%estimated ± 3.8 pp, low confidence
8Claude Opus 589.0%measured
9GPT-6 Astra89.0%estimated ± 1.9 pp, high confidence
10Sakana Fugu88.8%estimated ± 0.8 pp, low confidence
11Gemini 3.7 Flash88.7%measured
12Grok 4.788.6%estimated ± 2.7 pp, low confidence
13Step 5 Preview88.5%estimated ± 0.8 pp, low confidence
14Gemini 3.1 Pro88.5%measured
15Claude Mythos 588.5%estimated ± 4.7 pp, medium confidence
16Sakana Fugu-Ultra88.5%estimated ± 0.8 pp, low confidence
17Grok 4.688.2%measured
18Gemini 3.6 Flash88.1%measured
19SWE-288.1%estimated ± 1.9 pp, high confidence
20GPT-5.2-Codex88.0%measured
21Qwen3.8 Max87.9%measured
22Claude Opus 4.887.8%measured
23Gemini 3.5 Flash87.6%measured
24DeepSeek V4 Pro 081387.5%measured
25Grok 4.587.4%measured
26DeepSeek V4 Flash 073187.3%measured
27GPT-5.3 Codex87.3%measured
28Kimi K387.2%measured
29Qwen3.7 Max87.1%measured
30Claude Haiku 5.587.1%estimated ± 1.9 pp, high confidence
31Claude Sonnet 5.587.0%estimated ± 1.9 pp, high confidence
32Kimi K2.686.8%measured
33Muse Spark 1.386.8%estimated ± 2.7 pp, low confidence
34Hy4 preview86.7%estimated ± 4.7 pp, medium confidence
35Qwen3.7 Plus86.6%estimated ± 0.8 pp, medium confidence
36Ornith-1.5-397B86.5%estimated ± 4.7 pp, medium confidence
37Claude Opus 4.7 (Adaptive)86.3%estimated ± 4.7 pp, medium confidence
38Beam86.2%estimated ± 0.8 pp, medium confidence
39GPT-5.486.2%estimated ± 3.8 pp, low confidence
40Qwen3.8-Omni-Flash86.0%estimated ± 4.7 pp, medium confidence
41Nemotron 3 Ultra86.0%measured
42Qwen3.6 Plus86.0%measured
43Kimi K2.586.0%estimated ± 0.8 pp, medium confidence
44SWE-1.786.0%estimated ± 1.9 pp, high confidence
45GPT-5.6 Terra85.9%measured
46Inkling-Small85.9%measured
47Muse Spark 1.185.9%measured
48Qwen3.8-Flash-Next85.8%estimated ± 4.7 pp, medium confidence
49Ornith-1.0-397B85.7%estimated ± 4.7 pp, medium confidence
50Gemini 3 Flash85.6%measured
51GPT-5.1-Codex85.6%measured
52Qwen 3.6 Max (preview)85.6%estimated ± 0.8 pp, medium confidence
53Inkling85.5%measured
54GPT-5.585.3%measured
55dots3-note Preview85.2%estimated ± 4.7 pp, medium confidence
56Claude Opus 4.785.1%measured
57Muse Glimmer 30B84.7%estimated ± 0.8 pp, medium confidence
58Atria Dawn Preview84.7%estimated ± 4.7 pp, medium confidence
59Ornith-1.5-35B-A3B84.7%estimated ± 4.7 pp, medium confidence
60Laguna S 2.184.6%estimated ± 4.7 pp, medium confidence
61Grok 4.384.5%measured
62Grok 4.2084.3%measured
63Hy3 Preview84.1%estimated ± 0.8 pp, low confidence
64A.X K284.1%estimated ± 0.8 pp, low confidence
65GPT-5.6 Luna84.0%estimated ± 2.7 pp, low confidence
66GPT-5.4 nano84.0%measured
67Ling 3.0 Flash84.0%measured
68Qwen3.8-27B84.0%measured
69Ling 3.0 Flash FP883.9%estimated ± 0.8 pp, low confidence
70GPT-5.1-Codex-Max83.6%measured
71Granite 4.2 30B83.5%estimated ± 0.8 pp, low confidence
72Claude Opus 4.583.4%estimated ± 4.7 pp, medium confidence
73Mercury 2.583.3%estimated ± 0.8 pp, low confidence
74Qwen3.5 Flash83.3%measured
75K-EXAONE 2.083.2%estimated ± 0.8 pp, low confidence
76Step 3.7 Flash82.9%estimated ± 4.7 pp, medium confidence
77Granite 4.2 8B82.9%estimated ± 0.8 pp, low confidence
78GPT-5.6 Sol82.6%measured
79Gemma 4 31B82.6%estimated ± 3.8 pp, low confidence
80GPT-5.282.5%estimated ± 4.7 pp, medium confidence
81GLM-582.4%estimated ± 3.8 pp, low confidence
82Claude Sonnet 582.4%measured
83GLM-4.782.2%measured
84MiniMax M382.2%measured
85Claude Sonnet 4.682.1%measured
86Kimi K2.7 Code82.1%measured
87Nemotron 3 Nano Omni 30B A3B81.8%estimated ± 0.8 pp, low confidence
88Claude Opus 4.681.8%estimated ± 1.9 pp, high confidence
89Nemotron 3.5 Lightning 30B A3B NVFP481.7%estimated ± 0.8 pp, low confidence
90Grok 481.5%estimated ± 3.8 pp, low confidence
91GPT-5.4 mini81.5%measured
92MiMo-V2.581.5%measured
93GLM-5.181.4%measured
94MiMo-V2.5-Pro81.4%measured
95Agents-A1-4B81.2%estimated ± 0.8 pp, low confidence
96GPT-OSS 120B81.1%estimated ± 3.8 pp, low confidence
97DeepSeek V3.281.1%estimated ± 3.8 pp, low confidence
98GLM-4.681.0%measured
99GPT-OSS 20B80.9%estimated ± 3.8 pp, low confidence
100Qwen3.6-27B80.8%estimated ± 4.7 pp, medium confidence
101BTL-480.6%estimated ± 9.9 pp, low confidence
102Ling 2.6 Flash80.6%estimated ± 0.8 pp, low confidence
103GLM-5.380.5%measured
104GLM-5.3-Flash80.5%measured
105MiniCPM5-2B80.4%estimated ± 0.8 pp, low confidence
106MAI-Thinking-180.2%estimated ± 4.7 pp, medium confidence
107MiMo-V2-Pro80.1%estimated ± 9.9 pp, low confidence
108Gemini 3.1 Flash-Lite80.1%measured
109MiniMax M2.779.9%measured
110Granite 4.2 3B79.9%estimated ± 0.8 pp, low confidence
111Muse Spark79.8%estimated ± 4.7 pp, medium confidence
112Apodex 1.179.8%estimated ± 9.9 pp, low confidence
113Mistral Medium 3.5 128B79.6%estimated ± 9.9 pp, low confidence
114Claude Sonnet 4.579.1%estimated ± 9.9 pp, low confidence
115Gemini 3.5 Flash-Lite79.0%measured
116Composer 2.578.7%estimated ± 2.7 pp, low confidence
117Kimi K2.5 (Reasoning)78.6%estimated ± 9.9 pp, low confidence
118Qwen3.5 397B78.3%estimated ± 4.7 pp, medium confidence
119Ornith-1.0-35B77.7%estimated ± 4.7 pp, medium confidence
120DeepSeek V4.1 Flash77.4%estimated ± 6.1 pp, low confidence
121MiMo-V2.6-Flash77.4%estimated ± 6.1 pp, low confidence
122MiMo-V2.6-Pro77.4%estimated ± 6.1 pp, low confidence
123Qwen3.6-35B-A3B76.5%estimated ± 4.7 pp, medium confidence
124MiMo-V2-Omni75.8%estimated ± 9.9 pp, low confidence
125Claude 4.1 Opus75.3%estimated ± 9.9 pp, low confidence
126Laguna XS 2.173.7%estimated ± 4.7 pp, medium confidence
127MiMo-V2-Flash73.6%estimated ± 9.9 pp, low confidence
128Ornith-1.5-9B73.6%estimated ± 4.7 pp, medium confidence
129Claude 4 Sonnet72.5%estimated ± 9.9 pp, low confidence
130MAI-Code-1.1-Flash72.3%estimated ± 9.9 pp, low confidence
131Qwen3.5-27B72.0%estimated ± 9.9 pp, low confidence
132Qwen3.5-122B-A10B71.3%estimated ± 9.9 pp, low confidence
133GLM-5.269.5%measured
134Solar Pro 468.9%estimated ± 9.9 pp, low confidence
135Solar Open 268.5%estimated ± 9.9 pp, low confidence
136Laguna M.168.1%measured
137Laguna XS.267.8%measured
138GLM-4.567.4%measured
139Qwen3.5-35B-A3B66.2%estimated ± 9.9 pp, low confidence
140Ornith-1.0-9B64.1%estimated ± 4.7 pp, low confidence
141Grok Code Fast 162.0%measured
142LongCat-Flash-Lite-Sparse57.8%estimated ± 4.7 pp, low confidence
143Gemini 2.5 Pro54.6%estimated ± 9.9 pp, low confidence
144Ternary Bonsai 2 27B47.4%estimated ± 9.9 pp, low confidence
145Claude Haiku 4.541.2%measured
146GPT-4.132.0%estimated ± 9.9 pp, low confidence
147ZAYA1-74B-Preview28.7%estimated ± 9.9 pp, low confidence
148LLaDA2.2-flash20.8%estimated ± 4.7 pp, low confidence
149o3-mini20.3%estimated ± 9.9 pp, low confidence
150Claude 3.5 Sonnet19.8%estimated ± 9.9 pp, low confidence
151DeepSeek V38.9%estimated ± 9.9 pp, low confidence
152GPT-4.1 mini0.3%estimated ± 9.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General