benchgap
Agentic · tools

MCP-Tasks leaderboard

As of 2026-10-07, the highest measured score on MCP-Tasks is 74.2% by Qwen3.5 397B. 161 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude 3 Haiku100.0%estimated ± 8.2 pp, low confidence
2DeepSeek-R1100.0%estimated ± 8.2 pp, low confidence
3DeepSeek V3100.0%estimated ± 8.2 pp, low confidence
4DeepSeek V3.1100.0%estimated ± 8.2 pp, low confidence
5DeepSeek V3.1 (Reasoning)100.0%estimated ± 8.2 pp, low confidence
6Exaone 4.0 1.2B100.0%estimated ± 8.2 pp, low confidence
7Exaone 4.0 32B100.0%estimated ± 8.2 pp, low confidence
8Gemini 2.5 Flash100.0%estimated ± 8.2 pp, low confidence
9Gemma 3 27B100.0%estimated ± 8.2 pp, low confidence
10Gemma 4 12B100.0%estimated ± 8.2 pp, low confidence
11Gemma 4 E2B100.0%estimated ± 8.2 pp, low confidence
12Gemma 4 E4B100.0%estimated ± 8.2 pp, low confidence
13GPT-4.1 nano100.0%estimated ± 8.2 pp, low confidence
14GPT-4o100.0%estimated ± 8.2 pp, low confidence
15Granite-4.0-350M100.0%estimated ± 8.2 pp, low confidence
16Granite-4.0-H-1B100.0%estimated ± 8.2 pp, low confidence
17Granite-4.0-H-350M100.0%estimated ± 8.2 pp, low confidence
18K-EXAONE 2.0100.0%estimated ± 6.4 pp, low confidence
19LFM2.5-VL-1.6B-Extract100.0%estimated ± 8.2 pp, low confidence
20Llama 3.1 405B100.0%estimated ± 8.2 pp, low confidence
21Llama 4 Maverick100.0%estimated ± 8.2 pp, low confidence
22Llama 4 Scout100.0%estimated ± 8.2 pp, low confidence
23Mercury 2.5100.0%estimated ± 3.6 pp, low confidence
24Mistral Large 2100.0%estimated ± 8.2 pp, low confidence
25Mistral Large 3100.0%estimated ± 8.2 pp, low confidence
26Mistral Medium 3100.0%estimated ± 8.2 pp, low confidence
27Mistral Medium 3.5 128B100.0%estimated ± 3.6 pp, low confidence
28Mistral Small 4100.0%estimated ± 8.2 pp, low confidence
29Mistral Small 4 (Reasoning)100.0%estimated ± 8.2 pp, low confidence
30Nemotron 3 Nano 30B100.0%estimated ± 8.2 pp, low confidence
31Nemotron Ultra 253B100.0%estimated ± 8.2 pp, low confidence
32North Mini Code100.0%estimated ± 8.2 pp, low confidence
33Nova Pro100.0%estimated ± 8.2 pp, low confidence
34o3-mini100.0%estimated ± 8.2 pp, low confidence
35Ornith-1.0-397B100.0%estimated ± 6.4 pp, low confidence
36Phi-4100.0%estimated ± 8.2 pp, low confidence
37Qwen3-Omni-30B-A3B-Instruct100.0%estimated ± 8.2 pp, low confidence
38Qwen3-Omni-30B-A3B-Thinking100.0%estimated ± 8.2 pp, low confidence
39Sarvam 30B100.0%estimated ± 8.2 pp, low confidence
40Solar Pro 2100.0%estimated ± 8.2 pp, low confidence
41Ultravox v0.6 Llama 3.3 70B100.0%estimated ± 8.2 pp, low confidence
42Gemma 4 26B A4B99.0%estimated ± 8.2 pp, low confidence
43Qwen3.6-27B98.1%estimated ± 6.4 pp, low confidence
44Nemotron 3 Nano Omni 30B A3B97.9%estimated ± 8.2 pp, low confidence
45GLM-4.5-Air97.1%estimated ± 8.2 pp, low confidence
46Sarvam 105B96.9%estimated ± 8.2 pp, low confidence
47DeepSeek V3 032496.8%estimated ± 8.2 pp, low confidence
48GPT-4.196.8%estimated ± 8.2 pp, low confidence
49Claude Opus 4.693.7%estimated ± 6.4 pp, low confidence
50Claude 4 Sonnet93.5%estimated ± 8.2 pp, low confidence
51GPT-4.1 mini93.1%estimated ± 8.2 pp, low confidence
52Ornith-1.0-35B92.5%estimated ± 6.4 pp, low confidence
53Gemini 2.5 Pro92.3%estimated ± 8.2 pp, low confidence
54Gemma 4 31B88.6%estimated ± 8.2 pp, low confidence
55GPT-OSS 20B88.4%estimated ± 8.2 pp, low confidence
56Claude Sonnet 4.688.3%estimated ± 6.4 pp, low confidence
57Kimi K287.9%estimated ± 8.2 pp, low confidence
58o186.9%estimated ± 8.2 pp, low confidence
59Step 3.7 Flash86.9%estimated ± 6.4 pp, low confidence
60Grok 4.1 Fast86.2%estimated ± 8.2 pp, low confidence
61GPT-OSS 120B84.9%estimated ± 8.2 pp, low confidence
62Grok 4 Fast (Reasoning)84.9%estimated ± 8.2 pp, low confidence
63Qwen3.8 Max84.2%estimated ± 6.0 pp, low confidence
64Claude 4.1 Opus Thinking81.3%estimated ± 8.2 pp, low confidence
65GPT-5.6 Sol80.4%estimated ± 7.0 pp, low confidence
66Muse Spark80.4%estimated ± 6.4 pp, low confidence
67Claude Opus 4.779.7%estimated ± 8.2 pp, low confidence
68MiMo-V2.5-Pro79.7%estimated ± 3.6 pp, low confidence
69K-Exaone79.5%estimated ± 8.2 pp, low confidence
70Qwen3 Max79.5%estimated ± 8.2 pp, low confidence
71dots3-note Preview79.1%estimated ± 6.0 pp, low confidence
72Grok 479.1%estimated ± 8.2 pp, low confidence
73Ornith-1.0-9B79.0%estimated ± 6.4 pp, low confidence
74Grok Code Fast 178.6%estimated ± 8.2 pp, low confidence
75GPT-5.6 Luna78.3%estimated ± 7.0 pp, low confidence
76GPT-5.6 Terra78.2%estimated ± 7.0 pp, low confidence
77GLM-4.677.8%estimated ± 8.2 pp, low confidence
78MiMo-V2.577.5%estimated ± 6.4 pp, low confidence
79Muse Spark 1.177.1%estimated ± 0.6 pp, low confidence
80Claude Opus 577.1%estimated ± 0.6 pp, low confidence
81Step 5 Preview77.1%estimated ± 0.6 pp, low confidence
82Kimi K377.1%estimated ± 0.6 pp, low confidence
83Hy4 preview77.1%estimated ± 0.6 pp, low confidence
84Gemini 3.5 Flash77.1%estimated ± 0.6 pp, low confidence
85Claude Opus 4.877.1%estimated ± 0.6 pp, low confidence
86Ornith-1.5-397B77.1%estimated ± 0.6 pp, low confidence
87Inkling-Small77.1%estimated ± 0.6 pp, low confidence
88Beam77.1%estimated ± 0.6 pp, low confidence
89Claude Opus 4.7 (Adaptive)77.0%estimated ± 0.6 pp, low confidence
90GLM-5.277.0%estimated ± 0.6 pp, low confidence
91Qwen3.7 Max77.0%estimated ± 0.6 pp, low confidence
92Kimi K2.7 Code77.0%estimated ± 0.6 pp, low confidence
93Muse Glimmer 30B77.0%estimated ± 0.6 pp, low confidence
94Agents-A1-4B77.0%estimated ± 8.2 pp, low confidence
95GPT-5.577.0%estimated ± 0.6 pp, low confidence
96MiniMax M377.0%estimated ± 0.6 pp, low confidence
97Inkling77.0%estimated ± 0.6 pp, low confidence
98DeepSeek V4 Pro 081377.0%estimated ± 0.6 pp, low confidence
99Qwen3.7 Plus77.0%estimated ± 0.6 pp, low confidence
100GLM-5.176.9%estimated ± 0.6 pp, low confidence
101GPT-5.476.9%estimated ± 0.6 pp, low confidence
102Ornith-1.5-35B-A3B76.9%estimated ± 0.6 pp, low confidence
103DeepSeek V4 Flash 073176.9%estimated ± 0.6 pp, low confidence
104Ling 3.0 Flash76.7%estimated ± 0.6 pp, low confidence
105Qwen3.6-35B-A3B76.6%estimated ± 0.6 pp, low confidence
106Solar Pro 476.5%estimated ± 0.6 pp, low confidence
107Solar Open 276.3%estimated ± 0.6 pp, low confidence
108GPT-5.4 mini76.2%estimated ± 0.6 pp, low confidence
109GPT-5.4 nano76.0%estimated ± 0.6 pp, low confidence
110Kimi K2.676.0%estimated ± 0.6 pp, low confidence
111Ornith-1.5-9B75.8%estimated ± 0.6 pp, low confidence
112Ternary Bonsai 2 27B75.7%estimated ± 8.2 pp, low confidence
113o375.4%estimated ± 8.2 pp, low confidence
114Nemotron 3 Ultra75.1%estimated ± 3.6 pp, low confidence
115GPT-5.174.7%estimated ± 8.2 pp, low confidence
116Qwen3.5 397B74.2%measured
117Qwen3.6 Plus74.1%measured
118GPT-5.1-Codex74.0%estimated ± 8.2 pp, low confidence
119GPT-5.1-Codex-Max74.0%estimated ± 8.2 pp, low confidence
120LLaDA2.2-flash73.8%estimated ± 0.6 pp, medium confidence
121LongCat-Flash-Lite-Sparse73.5%estimated ± 0.6 pp, medium confidence
122MiMo-V2-Flash73.4%estimated ± 8.2 pp, low confidence
123Qwen3.5 397B (Reasoning)73.4%estimated ± 8.2 pp, low confidence
124GPT-5.272.8%estimated ± 8.2 pp, low confidence
125GPT-5 (high)72.8%estimated ± 8.2 pp, low confidence
126Command A+72.7%estimated ± 8.2 pp, low confidence
127GPT-5.3 Codex72.1%estimated ± 8.2 pp, low confidence
128Ling 2.6 Flash72.1%estimated ± 8.2 pp, low confidence
129Solar Pro 371.9%estimated ± 8.2 pp, low confidence
130Claude Opus 4.571.8%measured
131GPT-5 (medium)71.7%estimated ± 8.2 pp, low confidence
132Gemini 3 Pro71.4%estimated ± 8.2 pp, low confidence
133LFM2.5-8B-A1B70.7%estimated ± 8.2 pp, low confidence
134Qwen3.5-35B-A3B70.0%estimated ± 8.2 pp, low confidence
135Claude Opus 4.5 Thinking69.8%estimated ± 8.2 pp, low confidence
136Trinity-Large-Preview69.5%estimated ± 8.2 pp, low confidence
137Trinity-Large-Thinking69.5%estimated ± 8.2 pp, low confidence
138Gemini 3.1 Pro69.3%estimated ± 6.4 pp, low confidence
139MiMo-V2-Pro69.3%estimated ± 6.4 pp, low confidence
140Claude Opus 4.6 (Adaptive)68.2%estimated ± 8.2 pp, low confidence
141GPT-5.2-Codex68.2%estimated ± 8.2 pp, low confidence
142LLaDA2.2-mini68.1%estimated ± 6.4 pp, low confidence
143Grok 4.1 Fast (Reasoning)67.4%estimated ± 8.2 pp, low confidence
144Qwen3.5-122B-A10B67.2%estimated ± 8.2 pp, low confidence
145Qwen3.5-27B67.0%estimated ± 8.2 pp, low confidence
146GLM-4.765.8%estimated ± 8.2 pp, low confidence
147Kimi K2.5 (Reasoning)65.8%estimated ± 8.2 pp, low confidence
148Qwen 3.6 Max (preview)65.8%estimated ± 8.2 pp, low confidence
149GLM-5-Turbo65.8%estimated ± 6.4 pp, low confidence
150Grok 4.364.6%estimated ± 8.2 pp, low confidence
151Claude Fable 564.1%estimated ± 8.2 pp, low confidence
152Pokee-Isaac 28B63.1%estimated ± 3.6 pp, medium confidence
153GLM-5V-Turbo62.4%estimated ± 6.4 pp, low confidence
154GLM-560.8%measured
155Kimi K2.559.1%measured
156Gemini 3 Flash54.8%estimated ± 6.4 pp, low confidence
157MiniMax M2.754.0%estimated ± 6.4 pp, low confidence
158Granite 4.2 30B51.3%estimated ± 3.6 pp, low confidence
159MiMo-V2-Omni48.7%estimated ± 6.4 pp, low confidence
160DeepSeek V3.241.4%estimated ± 6.4 pp, low confidence
161Granite 4.2 8B40.2%estimated ± 3.6 pp, low confidence
162Granite 4.2 3B13.0%estimated ± 3.6 pp, low confidence
163Atria Dawn Preview7.3%estimated ± 3.6 pp, low confidence
164Nemotron 3 Super 100B0.8%estimated ± 6.4 pp, low confidence
165Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 3.6 pp, low confidence
166LFM2.5-2.6B0.0%estimated ± 3.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General