benchgap
Coding

LiveCodeBench Pro leaderboard

As of 2026-10-07, the highest measured score on LiveCodeBench Pro is 90.8% by Sakana Fugu-Ultra. 67 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.5100.0%estimated ± 6.0 pp, low confidence
2Claude Sonnet 5.598.0%estimated ± 6.0 pp, low confidence
3Claude Fable 5.197.9%estimated ± 6.0 pp, low confidence
4Claude Mythos 597.4%estimated ± 6.0 pp, low confidence
5Claude Fable 597.2%estimated ± 6.0 pp, low confidence
6Claude Opus 596.8%estimated ± 6.0 pp, low confidence
7Sakana Fugu-Ultra90.8%measured
8Claude Opus 4.890.4%estimated ± 6.0 pp, low confidence
9Qwen3.8 Max89.4%estimated ± 6.0 pp, low confidence
10Hy4 preview88.0%estimated ± 6.0 pp, low confidence
11Beam87.8%estimated ± 6.0 pp, low confidence
12Sakana Fugu87.8%measured
13Ornith-1.5-397B87.5%estimated ± 6.0 pp, low confidence
14GPT-5.487.5%measured
15Grok 4.587.3%estimated ± 6.0 pp, low confidence
16GPT-5.6 Sol87.2%estimated ± 6.0 pp, low confidence
17Claude Opus 4.7 (Adaptive)87.0%estimated ± 6.0 pp, low confidence
18GPT-5.6 Terra86.3%estimated ± 6.0 pp, low confidence
19Qwen3.8-Omni-Flash86.2%estimated ± 6.0 pp, low confidence
20Claude Sonnet 586.2%estimated ± 6.0 pp, low confidence
21GPT-5.6 Luna85.8%estimated ± 6.0 pp, low confidence
22Qwen3.8-Flash-Next85.6%estimated ± 6.0 pp, low confidence
23Ornith-1.0-397B85.4%estimated ± 6.0 pp, low confidence
24GLM-5.285.3%estimated ± 6.0 pp, low confidence
25Qwen3.8-27B85.0%estimated ± 6.0 pp, low confidence
26Muse Spark 1.184.9%estimated ± 6.0 pp, low confidence
27dots3-note Preview84.5%estimated ± 6.0 pp, low confidence
28Qwen3.7 Max84.2%estimated ± 6.0 pp, low confidence
29Atria Dawn Preview83.4%estimated ± 6.0 pp, low confidence
30Ornith-1.5-35B-A3B83.4%estimated ± 6.0 pp, low confidence
31Laguna S 2.183.3%estimated ± 6.0 pp, low confidence
32MiniMax M383.0%estimated ± 6.0 pp, low confidence
33Gemini 3.1 Pro82.9%measured
34GPT-5.582.6%estimated ± 6.0 pp, low confidence
35Kimi K2.682.6%estimated ± 6.0 pp, low confidence
36GLM-5.182.5%estimated ± 6.0 pp, low confidence
37Qwen3.7 Plus81.8%estimated ± 6.0 pp, low confidence
38Qwen 3.6 Max (preview)81.6%estimated ± 6.0 pp, low confidence
39MiMo-V2.5-Pro81.5%estimated ± 6.0 pp, low confidence
40Claude Opus 4.581.4%estimated ± 6.0 pp, low confidence
41GPT-5.3 Codex81.2%estimated ± 6.0 pp, low confidence
42Ling 3.0 Flash81.0%estimated ± 6.0 pp, low confidence
43Qwen3.6 Plus81.0%estimated ± 6.0 pp, low confidence
44Step 3.7 Flash80.8%estimated ± 6.0 pp, low confidence
45MiniMax M2.780.7%estimated ± 6.0 pp, low confidence
46MiMo-V2.580.6%estimated ± 6.0 pp, low confidence
47Inkling-Small80.5%estimated ± 6.0 pp, low confidence
48GPT-5.280.2%estimated ± 6.0 pp, low confidence
49DeepSeek V4 Pro 081380.0%estimated ± 6.0 pp, low confidence
50Muse Spark80.0%measured
51Gemini 3.5 Flash79.8%estimated ± 6.0 pp, low confidence
52GLM-579.8%estimated ± 6.0 pp, low confidence
53Inkling79.1%estimated ± 6.0 pp, low confidence
54Gemini 3.5 Flash-Lite79.0%estimated ± 6.0 pp, low confidence
55Qwen3.6-27B78.4%estimated ± 6.0 pp, low confidence
56MAI-Thinking-177.8%estimated ± 6.0 pp, low confidence
57DeepSeek V4 Flash 073177.7%estimated ± 6.0 pp, low confidence
58Muse Glimmer 30B76.4%estimated ± 6.0 pp, low confidence
59Qwen3.5 397B76.2%estimated ± 6.0 pp, low confidence
60Kimi K2.576.0%estimated ± 6.0 pp, low confidence
61Ornith-1.0-35B75.7%estimated ± 6.0 pp, low confidence
62Qwen3.6-35B-A3B74.9%estimated ± 6.0 pp, low confidence
63Laguna M.174.6%estimated ± 6.0 pp, low confidence
64Grok 4.2074.2%measured
65Laguna XS 2.173.2%estimated ± 6.0 pp, low confidence
66Ornith-1.5-9B73.1%estimated ± 6.0 pp, low confidence
67Laguna XS.271.9%estimated ± 6.0 pp, low confidence
68Claude Opus 4.670.7%measured
69Ornith-1.0-9B68.6%estimated ± 6.0 pp, low confidence
70LongCat-Flash-Lite-Sparse66.2%estimated ± 6.0 pp, low confidence
71Granite 4.2 30B57.9%estimated ± 6.0 pp, low confidence
72LLaDA2.2-flash54.0%estimated ± 6.0 pp, low confidence
73Granite 4.2 8B38.3%estimated ± 6.0 pp, low confidence
74MiniCPM5-2B30.4%estimated ± 6.0 pp, low confidence
75MiniCPM5-1B22.7%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General