benchgap
Coding

VulcanBench v3 leaderboard

As of 2026-10-07, the highest measured score on VulcanBench v3 is 89.9% by Grok 4.5. 52 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Grok 4.589.9%measured
2Claude Fable 589.5%measured
3DeepSeek V4 Flash 073188.4%measured
4GPT-6 Astra87.1%estimated ± 0.7 pp, low confidence
5Claude Fable 5.187.0%estimated ± 0.7 pp, low confidence
6Claude Opus 5.587.0%estimated ± 0.7 pp, low confidence
7Claude Sonnet 5.587.0%estimated ± 0.7 pp, low confidence
8Grok 4.787.0%estimated ± 0.7 pp, medium confidence
9Claude Opus 587.0%measured
10GPT-5.6 Sol87.0%measured
11GPT-5.6 Terra87.0%measured
12Grok 4.687.0%measured
13Muse Spark 1.287.0%measured
14Muse Spark 1.387.0%estimated ± 0.7 pp, medium confidence
15Gemini 3.8 Flash87.0%estimated ± 0.7 pp, medium confidence
16Gemini 3.7 Flash86.3%estimated ± 5.6 pp, low confidence
17Gemini 3.1 Pro86.2%estimated ± 5.6 pp, low confidence
18GPT-5.2-Codex85.8%estimated ± 5.6 pp, low confidence
19GPT-5.6 Luna85.5%measured
20DeepSeek V4 Pro 081385.4%estimated ± 5.6 pp, low confidence
21GPT-5.3 Codex85.2%estimated ± 5.6 pp, low confidence
22Qwen3.7 Max85.1%estimated ± 5.6 pp, low confidence
23Kimi K2.684.9%estimated ± 5.6 pp, low confidence
24Nemotron 3 Ultra84.3%estimated ± 5.6 pp, low confidence
25Qwen3.6 Plus84.3%estimated ± 5.6 pp, low confidence
26Inkling-Small84.2%estimated ± 5.6 pp, low confidence
27Muse Spark 1.184.2%estimated ± 5.6 pp, low confidence
28Gemini 3 Flash84.0%estimated ± 5.6 pp, low confidence
29GPT-5.1-Codex84.0%estimated ± 5.6 pp, low confidence
30Inkling83.9%estimated ± 5.6 pp, low confidence
31Claude Opus 4.783.7%estimated ± 5.6 pp, low confidence
32Grok 4.383.3%estimated ± 5.6 pp, low confidence
33Grok 4.2083.2%estimated ± 5.6 pp, low confidence
34GPT-5.4 nano83.0%estimated ± 5.6 pp, low confidence
35Ling 3.0 Flash83.0%estimated ± 5.6 pp, low confidence
36GPT-5.1-Codex-Max82.7%estimated ± 5.6 pp, low confidence
37Claude Opus 4.882.7%estimated ± 5.1 pp, low confidence
38Qwen3.8-27B82.6%measured
39Qwen3.5 Flash82.6%estimated ± 5.6 pp, low confidence
40GLM-4.782.0%estimated ± 5.6 pp, low confidence
41MiniMax M382.0%estimated ± 5.6 pp, low confidence
42Claude Sonnet 4.681.9%estimated ± 5.6 pp, low confidence
43GPT-5.4 mini81.6%estimated ± 5.6 pp, low confidence
44MiMo-V2.581.6%estimated ± 5.6 pp, low confidence
45GLM-5.181.6%estimated ± 5.6 pp, low confidence
46MiMo-V2.5-Pro81.6%estimated ± 5.6 pp, low confidence
47GLM-4.681.4%estimated ± 5.6 pp, low confidence
48Qwen3.8 Max81.2%measured
49GLM-5.3-Flash81.1%estimated ± 5.6 pp, low confidence
50Gemini 3.1 Flash-Lite80.9%estimated ± 5.6 pp, low confidence
51MiniMax M2.780.9%estimated ± 5.6 pp, low confidence
52Gemini 3.5 Flash-Lite80.5%estimated ± 5.6 pp, low confidence
53GPT-5.579.6%estimated ± 5.1 pp, low confidence
54GLM-5.378.3%measured
55Laguna M.177.6%estimated ± 5.6 pp, low confidence
56Laguna XS.277.6%estimated ± 5.6 pp, low confidence
57GLM-4.577.5%estimated ± 5.6 pp, low confidence
58Grok Code Fast 176.9%estimated ± 5.6 pp, low confidence
59GLM-5.276.7%estimated ± 5.1 pp, low confidence
60Claude Haiku 4.576.2%measured
61Gemini 3.6 Flash75.4%estimated ± 5.1 pp, low confidence
62Claude Sonnet 575.2%estimated ± 0.7 pp, low confidence
63Kimi K373.7%measured
64Kimi K2.7 Code72.0%estimated ± 5.1 pp, low confidence
65Gemini 3.5 Flash71.2%estimated ± 5.1 pp, low confidence
66Composer 2.570.0%estimated ± 0.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General