benchgap
Agentic · tools

DRACO leaderboard

As of 2026-10-07, the highest measured score on DRACO is 88.6% by Claude Opus 5. 117 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.589.6%estimated ± 3.0 pp, low confidence
2Claude Opus 588.6%measured
3Claude Fable 5.187.3%estimated ± 3.0 pp, medium confidence
4Claude Sonnet 5.587.0%measured
5Grok 4.786.5%estimated ± 3.0 pp, medium confidence
6MiMo-V2.6-Pro85.9%estimated ± 3.0 pp, medium confidence
7Muse Spark 1.385.8%estimated ± 3.0 pp, medium confidence
8Qwen3.8 Max Preview85.6%estimated ± 3.0 pp, medium confidence
9Ling 3.1 Flash85.5%measured
10GLM-5.385.2%estimated ± 3.0 pp, medium confidence
11Qwen3.8-Flash-Next84.8%estimated ± 3.0 pp, medium confidence
12Gemini 4 Argon84.6%estimated ± 3.0 pp, medium confidence
13Grok 4.684.6%estimated ± 3.0 pp, medium confidence
14Claude Fable 584.4%estimated ± 3.0 pp, medium confidence
15GPT-5.6 Sol84.4%estimated ± 3.0 pp, medium confidence
16MiMo-V2.6-Flash84.3%estimated ± 3.0 pp, medium confidence
17DeepSeek V4.1 Flash84.1%estimated ± 3.0 pp, medium confidence
18DeepSeek V4 Pro 081383.9%estimated ± 3.0 pp, medium confidence
19GPT-6.1 Sol83.6%estimated ± 3.0 pp, medium confidence
20Step 5 Preview83.3%measured
21GPT-6 Astra82.9%estimated ± 3.0 pp, medium confidence
22Kimi K382.8%estimated ± 3.0 pp, medium confidence
23GPT-6 Sol82.3%estimated ± 3.0 pp, medium confidence
24Muse Spark 1.281.6%estimated ± 3.0 pp, medium confidence
25Claude Sonnet 581.4%estimated ± 3.0 pp, medium confidence
26GPT-5.6 Luna81.3%estimated ± 3.0 pp, medium confidence
27Claude Opus 4.881.2%estimated ± 3.0 pp, medium confidence
28GPT-5.6 Terra81.1%estimated ± 3.0 pp, medium confidence
29DeepSeek V4 Flash 073180.8%estimated ± 3.0 pp, medium confidence
30GPT-6 Luna80.8%estimated ± 3.0 pp, medium confidence
31Gemini 3.8 Flash80.7%estimated ± 3.0 pp, medium confidence
32Mistral Large 480.5%estimated ± 3.0 pp, medium confidence
33Qwen3.8-27B80.5%estimated ± 3.0 pp, medium confidence
34Gemini 3.7 Flash79.8%estimated ± 3.0 pp, medium confidence
35Grok 4.579.8%estimated ± 3.0 pp, medium confidence
36GLM-5.279.5%estimated ± 3.0 pp, medium confidence
37Claude Opus 4.7 (Adaptive)79.1%estimated ± 3.0 pp, medium confidence
38GPT-5.579.1%estimated ± 3.0 pp, medium confidence
39Gemini 3.5 Flash78.9%estimated ± 3.0 pp, medium confidence
40Gemini 3.6 Flash77.7%estimated ± 3.0 pp, medium confidence
41Hy4 preview77.2%measured
42GPT-5.476.9%estimated ± 3.0 pp, medium confidence
43MiniMax M376.8%estimated ± 3.0 pp, medium confidence
44Hy3 Preview76.2%estimated ± 3.0 pp, medium confidence
45Muse Spark 1.176.2%estimated ± 3.0 pp, medium confidence
46Apodex 1.175.8%estimated ± 3.0 pp, medium confidence
47Apodex 1.1 Mini75.8%estimated ± 3.0 pp, medium confidence
48Quasar 438B75.6%estimated ± 3.0 pp, medium confidence
49Ling 3.0 Flash VL75.2%estimated ± 3.0 pp, medium confidence
50Nemotron 3 Ultra75.1%estimated ± 3.0 pp, medium confidence
51Qwen3.7 Max74.5%estimated ± 3.0 pp, medium confidence
52Inkling-Small74.3%estimated ± 3.0 pp, medium confidence
53MiMo-V2.5-Pro74.3%estimated ± 3.0 pp, medium confidence
54GLM-5.174.3%estimated ± 3.0 pp, medium confidence
55Solar Pro 474.1%estimated ± 3.0 pp, medium confidence
56Grok 4.373.5%estimated ± 3.0 pp, medium confidence
57Inkling73.4%estimated ± 3.0 pp, medium confidence
58Hy373.2%estimated ± 3.0 pp, medium confidence
59Kimi K2.672.6%estimated ± 3.0 pp, medium confidence
60Kimi K2.7 Code72.6%estimated ± 3.0 pp, medium confidence
61GLM-4.772.1%estimated ± 3.0 pp, medium confidence
62GPT-5.4 mini72.1%estimated ± 3.0 pp, medium confidence
63Step 3.7 Flash72.1%estimated ± 3.0 pp, medium confidence
64MiniMax M2.772.1%estimated ± 3.0 pp, medium confidence
65Muse Spark71.8%estimated ± 3.0 pp, medium confidence
66Qwen3.6 Plus71.7%estimated ± 3.0 pp, medium confidence
67Qwen3.6-27B71.6%estimated ± 3.0 pp, medium confidence
68Gemini 3.5 Flash-Lite71.5%estimated ± 3.0 pp, medium confidence
69A.X K271.1%estimated ± 3.0 pp, medium confidence
70GPT-5.4 nano70.8%estimated ± 3.0 pp, medium confidence
71Ling 3.0 Flash FP870.7%estimated ± 3.0 pp, medium confidence
72Ling 3.0 Flash70.4%measured
73GPT-5 (high)70.2%estimated ± 3.0 pp, low confidence
74Qwen3.6-35B-A3B69.7%estimated ± 3.0 pp, low confidence
75Kimi K2.568.6%estimated ± 3.0 pp, low confidence
76Kimi K2.5 (Reasoning)68.6%estimated ± 3.0 pp, low confidence
77GPT-5.168.3%estimated ± 3.0 pp, low confidence
78Qwen3.5-122B-A10B68.0%estimated ± 3.0 pp, low confidence
79Gemini 3.1 Pro67.6%estimated ± 3.0 pp, low confidence
80Muse Glimmer 30B67.5%estimated ± 3.0 pp, low confidence
81Qwen3.7 Plus67.1%estimated ± 3.0 pp, low confidence
82Mistral Medium 3.5 128B67.0%estimated ± 3.0 pp, low confidence
83MiniCPM5-2B65.9%estimated ± 3.0 pp, low confidence
84Nemotron 3.5 Lightning 30B A3B NVFP464.4%estimated ± 3.0 pp, low confidence
85MiMo-V2-Flash64.3%estimated ± 3.0 pp, low confidence
86Gemma 4 31B64.0%estimated ± 3.0 pp, low confidence
87GPT-OSS 120B63.8%estimated ± 3.0 pp, low confidence
88Granite 4.2 30B63.1%estimated ± 3.0 pp, low confidence
89Gemma 4 26B A4B62.9%estimated ± 3.0 pp, low confidence
90Ling 3.0 Tiny62.9%estimated ± 3.0 pp, low confidence
91Command A+61.9%estimated ± 3.0 pp, low confidence
92Celeris-161.5%estimated ± 3.0 pp, low confidence
93DeepSeek V361.5%estimated ± 3.0 pp, low confidence
94DeepSeek V3 032461.5%estimated ± 3.0 pp, low confidence
95Gemini 2.5 Pro61.5%estimated ± 3.0 pp, low confidence
96Gemma 3 27B61.5%estimated ± 3.0 pp, low confidence
97Gemma 4 12B61.5%estimated ± 3.0 pp, low confidence
98Gemma 4 E2B61.5%estimated ± 3.0 pp, low confidence
99Gemma 4 E4B61.5%estimated ± 3.0 pp, low confidence
100GPT-4.1 mini61.5%estimated ± 3.0 pp, low confidence
101GPT-4.1 nano61.5%estimated ± 3.0 pp, low confidence
102GPT-4o61.5%estimated ± 3.0 pp, low confidence
103GPT-4o mini61.5%estimated ± 3.0 pp, low confidence
104GPT-OSS 20B61.5%estimated ± 3.0 pp, low confidence
105Granite 4.2 3B61.5%estimated ± 3.0 pp, low confidence
106Granite 4.2 8B61.5%estimated ± 3.0 pp, low confidence
107K-Exaone61.5%estimated ± 3.0 pp, low confidence
108LFM2.5-2.6B61.5%estimated ± 3.0 pp, low confidence
109Ling 2.6 Flash61.5%estimated ± 3.0 pp, low confidence
110Llama 4 Maverick61.5%estimated ± 3.0 pp, low confidence
111Llama 4 Scout61.5%estimated ± 3.0 pp, low confidence
112Mercury 2.561.5%estimated ± 3.0 pp, low confidence
113Mistral Large 361.5%estimated ± 3.0 pp, low confidence
114Mistral Small 461.5%estimated ± 3.0 pp, low confidence
115Mistral Small 4 (Reasoning)61.5%estimated ± 3.0 pp, low confidence
116Nemotron 3 Nano 30B61.5%estimated ± 3.0 pp, low confidence
117Nemotron 3 Nano Omni 30B A3B61.5%estimated ± 3.0 pp, low confidence
118Nemotron 3 Super 100B61.5%estimated ± 3.0 pp, low confidence
119North Mini Code61.5%estimated ± 3.0 pp, low confidence
120Solar Pro 361.5%estimated ± 3.0 pp, low confidence
121Trinity-Large-Preview61.5%estimated ± 3.0 pp, low confidence
122Trinity-Large-Thinking61.5%estimated ± 3.0 pp, low confidence
123Ultravox v0.6 Llama 3.3 70B61.5%estimated ± 3.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General