benchgap
Agentic · tools

MCP Atlas leaderboard

As of 2026-10-07, the highest measured score on MCP Atlas is 88.1% by Muse Spark 1.1. 151 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Grok 4.6100.0%estimated ± 0.7 pp, low confidence
2Gemini 4 Argon89.2%estimated ± 2.7 pp, low confidence
3Muse Spark 1.188.1%measured
4Claude Sonnet 5.588.0%estimated ± 2.7 pp, low confidence
5Claude Opus 5.587.5%estimated ± 2.7 pp, low confidence
6DeepSeek V4.1 Flash87.4%estimated ± 2.7 pp, low confidence
7GPT-6 Astra87.3%estimated ± 2.7 pp, low confidence
8Grok 4.786.7%estimated ± 2.7 pp, low confidence
9GPT-6.1 Sol86.5%estimated ± 2.7 pp, low confidence
10GPT-5.6 Sol86.3%estimated ± 3.5 pp, low confidence
11GLM-5.385.9%estimated ± 2.7 pp, low confidence
12GPT-6 Sol85.8%estimated ± 2.7 pp, low confidence
13Claude Opus 585.8%measured
14Step 5 Preview85.6%measured
15Apodex 1.185.6%estimated ± 0.7 pp, low confidence
16GLM-5.3-Flash85.6%estimated ± 2.7 pp, low confidence
17Gemini 3.8 Flash85.5%estimated ± 2.7 pp, low confidence
18Mistral Large 485.5%estimated ± 2.7 pp, low confidence
19Claude Fable 5.185.3%estimated ± 2.7 pp, low confidence
20MiMo-V2.6-Pro85.2%estimated ± 2.7 pp, low confidence
21Muse Spark 1.385.0%estimated ± 2.7 pp, medium confidence
22Gemini 3.7 Flash84.8%estimated ± 6.2 pp, low confidence
23Kimi K384.2%measured
24Ling 3.1 Flash84.2%estimated ± 11.1 pp, low confidence
25GPT-6 Luna84.0%estimated ± 2.7 pp, medium confidence
26Hy4 preview83.7%measured
27Gemini 3.5 Flash83.6%measured
28Qwen3.8-27B82.9%estimated ± 2.7 pp, medium confidence
29GPT-5.6 Terra82.6%estimated ± 3.5 pp, medium confidence
30Claude Opus 4.882.2%measured
31Claude Fable 582.1%estimated ± 5.5 pp, low confidence
32Holo3-35B-A3B81.3%estimated ± 11.5 pp, low confidence
33MiMo-V2.6-Flash81.0%estimated ± 7.8 pp, medium confidence
34Qwen3.8-Flash-Next80.9%estimated ± 7.8 pp, medium confidence
35Claude Haiku 5.580.2%estimated ± 2.7 pp, medium confidence
36Ornith-1.5-397B80.0%measured
37Claude Opus 4.779.6%estimated ± 10.4 pp, low confidence
38Inkling-Small79.6%measured
39Claude Sonnet 579.5%estimated ± 7.3 pp, medium confidence
40Fugu Cyber79.0%estimated ± 11.4 pp, low confidence
41Gemini 3.8 Flash Cyber78.9%estimated ± 11.4 pp, low confidence
42Qwen3.8 Max78.9%estimated ± 7.3 pp, medium confidence
43Beam78.7%measured
44GPT-5.6 Luna78.7%estimated ± 7.8 pp, low confidence
45Grok 4.578.7%estimated ± 7.8 pp, low confidence
46Muse Spark 1.278.7%estimated ± 7.8 pp, low confidence
47Quasar 438B78.7%estimated ± 7.8 pp, low confidence
48Qwen3.8 Max Preview78.7%estimated ± 7.8 pp, low confidence
49Gemini 3.6 Flash78.7%estimated ± 7.8 pp, low confidence
50Gemini 3.5 Flash Cyber78.5%estimated ± 11.4 pp, low confidence
51Claude Mythos Preview78.5%estimated ± 11.4 pp, low confidence
52Hy378.3%estimated ± 7.8 pp, low confidence
53Hy3 Preview78.3%estimated ± 7.8 pp, low confidence
54GPT-5.5 Pro78.2%estimated ± 10.0 pp, low confidence
55Atria Dawn Preview78.1%estimated ± 8.2 pp, medium confidence
56GPT-5.4 Pro77.8%estimated ± 10.0 pp, low confidence
57K-EXAONE 2.077.5%estimated ± 11.2 pp, low confidence
58Claude Opus 4.7 (Adaptive)77.3%measured
59Claude Mythos 577.1%estimated ± 10.0 pp, low confidence
60dots3-note Preview77.0%estimated ± 0.7 pp, medium confidence
61Ornith-1.0-397B77.0%estimated ± 11.2 pp, low confidence
62GLM-5.276.8%measured
63Holo3-122B-A10B76.6%estimated ± 11.5 pp, low confidence
64Qwen3.7 Max76.4%measured
65Kimi K2.7 Code76.0%measured
66Muse Glimmer 30B75.5%measured
67GPT-5.575.3%measured
68Ling 3.0 Flash Fin75.1%estimated ± 0.7 pp, medium confidence
69UI-Mate-27B74.3%estimated ± 11.5 pp, low confidence
70MiniMax M374.2%measured
71Inkling74.1%measured
72Agents-A173.8%estimated ± 7.3 pp, medium confidence
73DeepSeek V4 Pro 081373.6%measured
74Step 3.7 Flash73.5%estimated ± 7.3 pp, medium confidence
75Apodex 1.1 Mini73.3%estimated ± 0.7 pp, medium confidence
76Nemotron 3 Ultra73.3%estimated ± 2.7 pp, low confidence
77Qwen3.7 Plus73.2%measured
78Claude Opus 4.6 (Adaptive)72.3%estimated ± 9.4 pp, medium confidence
79GLM-5.171.8%measured
80Qwen 3.6 Max (preview)71.1%estimated ± 12.3 pp, low confidence
81GPT-5.470.6%measured
82Ling 3.0 Flash VL70.3%estimated ± 11.1 pp, low confidence
83Ornith-1.5-35B-A3B70.2%measured
84Ornith-1.0-35B69.3%estimated ± 11.2 pp, low confidence
85DeepSeek V4 Flash 073169.0%measured
86MiMo-V2.5-Pro68.5%estimated ± 7.8 pp, low confidence
87Laguna S 2.167.4%estimated ± 7.8 pp, medium confidence
88Ling 3.0 Flash65.5%measured
89Muse Spark65.4%estimated ± 11.1 pp, low confidence
90Ling 3.0 Flash FP865.0%estimated ± 7.8 pp, low confidence
91Qwen3.6-27B64.7%estimated ± 7.8 pp, low confidence
92Grok 4.364.6%estimated ± 7.8 pp, low confidence
93MiniMax M2.764.6%estimated ± 7.8 pp, low confidence
94Gemini 3.5 Flash-Lite64.6%estimated ± 7.8 pp, low confidence
95Celeris-164.6%estimated ± 7.8 pp, low confidence
96Command A+64.6%estimated ± 7.8 pp, low confidence
97DeepSeek V364.6%estimated ± 7.8 pp, low confidence
98Gemini 2.5 Pro64.6%estimated ± 7.8 pp, low confidence
99Gemini 3.1 Pro64.6%estimated ± 7.8 pp, low confidence
100Gemma 3 27B64.6%estimated ± 7.8 pp, low confidence
101Gemma 4 31B64.6%estimated ± 7.8 pp, low confidence
102GPT-OSS 120B64.6%estimated ± 7.8 pp, low confidence
103GPT-OSS 20B64.6%estimated ± 7.8 pp, low confidence
104Granite 4.2 8B64.6%estimated ± 7.8 pp, low confidence
105Llama 4 Maverick64.6%estimated ± 7.8 pp, low confidence
106Llama 4 Scout64.6%estimated ± 7.8 pp, low confidence
107Mistral Large 364.6%estimated ± 7.8 pp, low confidence
108Mistral Medium 3.5 128B64.6%estimated ± 7.8 pp, low confidence
109Mistral Small 464.6%estimated ± 7.8 pp, low confidence
110Mistral Small 4 (Reasoning)64.6%estimated ± 7.8 pp, low confidence
111Nemotron 3.5 Lightning 30B A3B NVFP464.6%estimated ± 7.8 pp, low confidence
112Nemotron 3 Nano 30B64.6%estimated ± 7.8 pp, low confidence
113Nemotron 3 Super 100B64.6%estimated ± 7.8 pp, low confidence
114Qwen3.5-122B-A10B64.6%estimated ± 7.8 pp, low confidence
115Trinity-Large-Preview64.6%estimated ± 7.8 pp, low confidence
116Trinity-Large-Thinking64.6%estimated ± 7.8 pp, low confidence
117Agents-A1-4B64.5%estimated ± 10.0 pp, low confidence
118A.X K264.2%estimated ± 11.1 pp, low confidence
119Qwen3.6-35B-A3B62.8%measured
120UI-Mate-9B62.4%estimated ± 11.5 pp, low confidence
121Claude Sonnet 4.662.2%estimated ± 8.2 pp, medium confidence
122Claude Opus 4.661.9%estimated ± 8.2 pp, medium confidence
123Solar Pro 461.4%measured
124Qwen3.5-27B60.6%estimated ± 10.0 pp, low confidence
125Qwen3.5-35B-A3B60.6%estimated ± 10.0 pp, low confidence
126GPT-5.160.2%estimated ± 11.1 pp, low confidence
127GPT-5.259.1%estimated ± 8.2 pp, medium confidence
128Ornith-1.0-9B58.5%estimated ± 11.2 pp, low confidence
129GPT-5.3 Codex58.4%estimated ± 8.2 pp, medium confidence
130Solar Open 258.2%measured
131GPT-5.4 mini57.7%measured
132MiMo-V2.557.0%estimated ± 11.2 pp, low confidence
133MiniCPM5-2B56.5%estimated ± 11.1 pp, low confidence
134GPT-5.4 nano56.1%measured
135Kimi K2.655.9%measured
136Ornith-1.5-9B54.2%measured
137MiMo-V2-Flash54.2%estimated ± 11.1 pp, low confidence
138GLM-4.754.1%estimated ± 10.0 pp, low confidence
139Granite 4.2 30B52.5%estimated ± 11.1 pp, low confidence
140Gemma 4 26B A4B52.2%estimated ± 11.1 pp, low confidence
141Ling 3.0 Tiny52.1%estimated ± 11.1 pp, low confidence
142Claude Sonnet 4.551.2%estimated ± 8.2 pp, medium confidence
143DeepSeek V3 032450.2%estimated ± 11.1 pp, low confidence
144Gemma 4 12B50.2%estimated ± 11.1 pp, low confidence
145Gemma 4 E2B50.2%estimated ± 11.1 pp, low confidence
146Gemma 4 E4B50.2%estimated ± 11.1 pp, low confidence
147GPT-4.1 mini50.2%estimated ± 11.1 pp, low confidence
148GPT-4.1 nano50.2%estimated ± 11.1 pp, low confidence
149GPT-4o50.2%estimated ± 11.1 pp, low confidence
150GPT-4o mini50.2%estimated ± 11.1 pp, low confidence
151Granite 4.2 3B50.2%estimated ± 11.1 pp, low confidence
152K-Exaone50.2%estimated ± 11.1 pp, low confidence
153LFM2.5-2.6B50.2%estimated ± 11.1 pp, low confidence
154Ling 2.6 Flash50.2%estimated ± 11.1 pp, low confidence
155Mercury 2.550.2%estimated ± 11.1 pp, low confidence
156Nemotron 3 Nano Omni 30B A3B50.2%estimated ± 11.1 pp, low confidence
157North Mini Code50.2%estimated ± 11.1 pp, low confidence
158Solar Pro 350.2%estimated ± 11.1 pp, low confidence
159Ultravox v0.6 Llama 3.3 70B50.2%estimated ± 11.1 pp, low confidence
160GPT-5.1-Codex49.4%estimated ± 8.2 pp, medium confidence
161GPT-5.2-Codex49.2%estimated ± 8.2 pp, medium confidence
162Qwen3.6 Plus48.2%measured
163MiMo-V2-Pro47.3%estimated ± 11.2 pp, low confidence
164LLaDA2.2-flash46.2%measured
165Qwen3.5 397B46.1%measured
166LLaDA2.2-mini45.8%estimated ± 11.2 pp, low confidence
167LongCat-Flash-Lite-Sparse45.6%measured
168Claude 4.1 Opus44.3%estimated ± 8.2 pp, medium confidence
169Gemini 3.1 Flash-Lite43.5%estimated ± 14.8 pp, low confidence
170GPT-4.143.5%estimated ± 14.8 pp, low confidence
171Grok 443.5%estimated ± 14.8 pp, low confidence
172Grok 4.1 Fast43.5%estimated ± 14.8 pp, low confidence
173Grok 4.2043.5%estimated ± 14.8 pp, low confidence
174Grok Build 0.143.5%estimated ± 14.8 pp, low confidence
175Qwen3 Max43.5%estimated ± 14.8 pp, low confidence
176GLM-5-Turbo42.6%estimated ± 11.2 pp, low confidence
177Claude Opus 4.542.3%measured
178Qwen3.5 Plus40.3%estimated ± 8.2 pp, medium confidence
179Claude 4 Sonnet40.1%estimated ± 8.2 pp, medium confidence
180Grok 4.138.6%estimated ± 13.5 pp, low confidence
181GLM-5V-Turbo37.8%estimated ± 11.2 pp, low confidence
182Claude Haiku 4.537.3%estimated ± 8.2 pp, medium confidence
183Kimi K2.5 (Reasoning)33.9%estimated ± 9.4 pp, medium confidence
184Gemini 3 Flash31.8%estimated ± 8.2 pp, medium confidence
185Gemini 3 Pro31.8%estimated ± 8.2 pp, medium confidence
186GLM-531.1%measured
187Kimi K2.529.5%measured
188GPT-5 (high)28.4%estimated ± 8.2 pp, low confidence
189MiMo-V2-Omni18.3%estimated ± 11.2 pp, low confidence
190DeepSeek V3.210.1%estimated ± 11.2 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General