benchgap
Agentic · tools

DeepPlanning leaderboard

As of 2026-10-07, the highest measured score on DeepPlanning is 62.3% by Qwen3.7 Plus. 50 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.7 Max100.0%estimated ± 9.1 pp, low confidence
2Qwen3.7 Plus62.3%measured
3Muse Spark 1.153.8%estimated ± 13.1 pp, low confidence
4Claude Opus 552.7%estimated ± 13.1 pp, low confidence
5Step 5 Preview52.6%estimated ± 13.1 pp, low confidence
6Kimi K352.0%estimated ± 13.1 pp, low confidence
7Hy4 preview51.8%estimated ± 13.1 pp, low confidence
8Gemini 3.5 Flash51.7%estimated ± 13.1 pp, low confidence
9Claude Opus 4.851.1%estimated ± 13.1 pp, low confidence
10Ornith-1.5-397B50.1%estimated ± 13.1 pp, low confidence
11Inkling-Small49.9%estimated ± 13.1 pp, low confidence
12Beam49.4%estimated ± 13.1 pp, low confidence
13Claude Opus 4.7 (Adaptive)48.8%estimated ± 13.1 pp, low confidence
14GLM-5.248.5%estimated ± 13.1 pp, low confidence
15Kimi K2.7 Code48.2%estimated ± 13.1 pp, low confidence
16Muse Glimmer 30B47.9%estimated ± 13.1 pp, low confidence
17GPT-5.547.8%estimated ± 13.1 pp, low confidence
18MiniMax M347.3%estimated ± 13.1 pp, low confidence
19Inkling47.2%estimated ± 13.1 pp, low confidence
20DeepSeek V4 Pro 081347.0%estimated ± 13.1 pp, low confidence
21GPT-5.445.5%estimated ± 13.1 pp, low confidence
22Ornith-1.5-35B-A3B45.3%estimated ± 13.1 pp, low confidence
23DeepSeek V4 Flash 073144.7%estimated ± 13.1 pp, low confidence
24Qwen 3.6 Max (preview)43.2%estimated ± 11.8 pp, low confidence
25Ling 3.0 Flash42.9%estimated ± 13.1 pp, low confidence
26Qwen3.6 Plus41.5%measured
27Solar Pro 440.8%estimated ± 13.1 pp, low confidence
28Solar Open 239.1%estimated ± 13.1 pp, low confidence
29GPT-5.4 mini38.8%estimated ± 13.1 pp, low confidence
30GPT-5.4 nano37.9%estimated ± 13.1 pp, low confidence
31Kimi K2.637.8%estimated ± 13.1 pp, low confidence
32Qwen3.5 397B37.6%measured
33Ornith-1.5-9B36.9%estimated ± 13.1 pp, low confidence
34Mercury 2.535.4%estimated ± 8.3 pp, low confidence
35Mistral Medium 3.5 128B35.4%estimated ± 8.3 pp, low confidence
36MiMo-V2.5-Pro35.4%estimated ± 8.3 pp, low confidence
37Nemotron 3 Ultra35.4%estimated ± 8.3 pp, low confidence
38GLM-5.135.4%estimated ± 8.3 pp, low confidence
39LLaDA2.2-flash32.3%estimated ± 13.1 pp, low confidence
40Agents-A1-4B26.9%estimated ± 9.1 pp, low confidence
41Agents-A126.4%estimated ± 9.1 pp, low confidence
42Claude Opus 4.526.4%measured
43Claude Sonnet 4.526.2%estimated ± 9.1 pp, low confidence
44DeepSeek V3.226.2%estimated ± 9.1 pp, low confidence
45GLM-4.726.2%estimated ± 9.1 pp, low confidence
46LongCat-Flash-Lite-Sparse26.2%estimated ± 9.1 pp, low confidence
47Qwen3.6-35B-A3B25.9%measured
48Qwen3.6-27B25.3%estimated ± 11.8 pp, low confidence
49Pokee-Isaac 28B16.9%estimated ± 8.3 pp, low confidence
50GLM-514.6%measured
51Kimi K2.514.4%measured
52Granite 4.2 30B13.0%estimated ± 8.3 pp, low confidence
53Atria Dawn Preview13.0%estimated ± 8.3 pp, low confidence
54Granite 4.2 3B13.0%estimated ± 8.3 pp, low confidence
55Granite 4.2 8B13.0%estimated ± 8.3 pp, low confidence
56LFM2.5-2.6B13.0%estimated ± 8.3 pp, low confidence
57Nemotron 3.5 Lightning 30B A3B NVFP413.0%estimated ± 8.3 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General