benchgap
Knowledge & reasoning

HealthBench Professional leaderboard

As of 2026-10-07, the highest measured score on HealthBench Professional is 69.2% by Claude Sonnet 5.5. 226 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.569.2%measured
2Claude Opus 5.565.6%measured
3Ling 3.1 Flash65.4%measured
4Gemini 4 Argon64.9%estimated ± 4.8 pp, medium confidence
5GPT-6 Astra64.7%measured
6GPT-6.1 Sol64.2%measured
7MiMo-V2.6-Pro61.1%estimated ± 4.8 pp, medium confidence
8GPT-6 Luna60.8%measured
9GPT-6 Sol60.8%measured
10Sakana Fugu60.7%estimated ± 4.8 pp, medium confidence
11Sakana Fugu-Ultra60.7%estimated ± 4.8 pp, medium confidence
12GPT-5.6 Sol60.5%measured
13Muse Spark 1.360.2%estimated ± 4.8 pp, medium confidence
14Claude Fable 5.160.0%estimated ± 1.3 pp, low confidence
15Claude Fable 560.0%estimated ± 1.3 pp, medium confidence
16Gemini 3.1 Pro60.0%estimated ± 1.3 pp, medium confidence
17Gemini 3.8 Flash60.0%estimated ± 1.3 pp, medium confidence
18Gemini 3.7 Flash60.0%estimated ± 1.3 pp, medium confidence
19Claude Opus 4.759.9%estimated ± 1.3 pp, medium confidence
20Claude Opus 4.859.9%estimated ± 1.3 pp, medium confidence
21Gemini 3.5 Flash59.9%estimated ± 1.3 pp, medium confidence
22Grok 4.659.9%estimated ± 1.3 pp, medium confidence
23Gemini 3.6 Flash59.9%estimated ± 1.3 pp, medium confidence
24Qwen3.7 Max59.9%estimated ± 1.3 pp, medium confidence
25Grok 4.559.9%estimated ± 1.3 pp, medium confidence
26Claude Opus 559.8%measured
27Gemini 3 Flash59.7%estimated ± 1.3 pp, medium confidence
28Qwen3.8 Max59.7%estimated ± 1.3 pp, medium confidence
29Muse Spark 1.259.6%estimated ± 1.3 pp, medium confidence
30GPT-5.559.5%estimated ± 1.3 pp, medium confidence
31Kimi K359.5%estimated ± 1.3 pp, medium confidence
32Muse Spark 1.159.3%measured
33Qwen3.6 Plus59.2%estimated ± 1.3 pp, medium confidence
34Kimi K2.659.1%estimated ± 1.3 pp, medium confidence
35Claude Sonnet 559.0%estimated ± 1.3 pp, medium confidence
36Claude Sonnet 4.658.8%estimated ± 1.3 pp, medium confidence
37Muse Spark58.8%estimated ± 1.3 pp, medium confidence
38DeepSeek V4 Pro 081358.3%estimated ± 1.3 pp, medium confidence
39GLM-5.158.1%estimated ± 1.3 pp, medium confidence
40Claude Mythos 558.1%estimated ± 4.8 pp, medium confidence
41GLM-5.357.9%estimated ± 1.3 pp, medium confidence
42Qwen3.8 Max Preview57.7%estimated ± 4.8 pp, medium confidence
43GPT-5.3 Codex57.7%estimated ± 4.8 pp, medium confidence
44GPT-5.6 Terra57.7%measured
45GLM-5.257.7%estimated ± 1.3 pp, medium confidence
46GPT-5.2-Codex57.7%estimated ± 4.8 pp, low confidence
47Kimi K2.7 Code57.7%estimated ± 4.8 pp, low confidence
48MiMo-V2.6-Flash57.7%estimated ± 4.8 pp, low confidence
49Mistral Large 457.7%estimated ± 4.8 pp, low confidence
50Apodex 1.157.7%estimated ± 4.8 pp, low confidence
51Apodex 1.1 Mini57.7%estimated ± 4.8 pp, low confidence
52Hy357.7%estimated ± 4.8 pp, low confidence
53Celeris-157.7%estimated ± 4.8 pp, low confidence
54Claude 3 Haiku57.7%estimated ± 4.8 pp, low confidence
55Claude 3 Opus57.7%estimated ± 4.8 pp, low confidence
56Claude 4.1 Opus Thinking57.7%estimated ± 4.8 pp, low confidence
57Claude 4 Sonnet57.7%estimated ± 4.8 pp, low confidence
58Command A+57.7%estimated ± 4.8 pp, low confidence
59DeepSeek-R157.7%estimated ± 4.8 pp, low confidence
60DeepSeek R1 Distill Qwen 32B57.7%estimated ± 4.8 pp, low confidence
61DeepSeek V3 032457.7%estimated ± 4.8 pp, low confidence
62DeepSeek V3.157.7%estimated ± 4.8 pp, low confidence
63DeepSeek V3.1 (Reasoning)57.7%estimated ± 4.8 pp, low confidence
64DeepSeek V3.257.7%estimated ± 4.8 pp, low confidence
65Exaone 4.0 1.2B57.7%estimated ± 4.8 pp, low confidence
66Exaone 4.0 32B57.7%estimated ± 4.8 pp, low confidence
67Gemini 1.0 Pro57.7%estimated ± 4.8 pp, low confidence
68Gemini 1.5 Pro57.7%estimated ± 4.8 pp, low confidence
69Gemini 2.5 Flash57.7%estimated ± 4.8 pp, low confidence
70Gemma 3 27B57.7%estimated ± 4.8 pp, low confidence
71GLM-4.5-Air57.7%estimated ± 4.8 pp, low confidence
72GLM-5-Turbo57.7%estimated ± 4.8 pp, low confidence
73GLM-5V-Turbo57.7%estimated ± 4.8 pp, low confidence
74GPT-4 Turbo57.7%estimated ± 4.8 pp, low confidence
75GPT-4o57.7%estimated ± 4.8 pp, low confidence
76GPT-4o mini57.7%estimated ± 4.8 pp, low confidence
77GPT-5.157.7%estimated ± 4.8 pp, low confidence
78GPT-5.1-Codex57.7%estimated ± 4.8 pp, low confidence
79GPT-5.1-Codex-Max57.7%estimated ± 4.8 pp, low confidence
80GPT-5 (high)57.7%estimated ± 4.8 pp, low confidence
81GPT-5 (medium)57.7%estimated ± 4.8 pp, low confidence
82GPT-OSS 120B57.7%estimated ± 4.8 pp, low confidence
83GPT-OSS 20B57.7%estimated ± 4.8 pp, low confidence
84Granite-4.0-350M57.7%estimated ± 4.8 pp, low confidence
85Granite-4.0-H-1B57.7%estimated ± 4.8 pp, low confidence
86Granite-4.0-H-350M57.7%estimated ± 4.8 pp, low confidence
87Grok 457.7%estimated ± 4.8 pp, low confidence
88Grok 4.1 Fast57.7%estimated ± 4.8 pp, low confidence
89Grok 4.1 Fast (Reasoning)57.7%estimated ± 4.8 pp, low confidence
90Grok 4 Fast (Reasoning)57.7%estimated ± 4.8 pp, low confidence
91Grok Code Fast 157.7%estimated ± 4.8 pp, low confidence
92K-Exaone57.7%estimated ± 4.8 pp, low confidence
93Kimi K257.7%estimated ± 4.8 pp, low confidence
94LFM2.5-2.6B57.7%estimated ± 4.8 pp, low confidence
95LFM2.5-8B-A1B57.7%estimated ± 4.8 pp, low confidence
96LFM2.5-VL-1.6B-Extract57.7%estimated ± 4.8 pp, low confidence
97Ling 3.0 Flash VL57.7%estimated ± 4.8 pp, low confidence
98Ling 3.0 Tiny57.7%estimated ± 4.8 pp, low confidence
99Llama 3.1 405B57.7%estimated ± 4.8 pp, low confidence
100Llama 4 Maverick57.7%estimated ± 4.8 pp, low confidence
101Llama 4 Scout57.7%estimated ± 4.8 pp, low confidence
102MiMo-V2-Omni57.7%estimated ± 4.8 pp, low confidence
103MiMo-V2-Pro57.7%estimated ± 4.8 pp, low confidence
104Mistral Large 257.7%estimated ± 4.8 pp, low confidence
105Mistral Large 357.7%estimated ± 4.8 pp, low confidence
106Mistral Medium 357.7%estimated ± 4.8 pp, low confidence
107Mistral Small 457.7%estimated ± 4.8 pp, low confidence
108Mistral Small 4 (Reasoning)57.7%estimated ± 4.8 pp, low confidence
109Muse Glimmer 30B57.7%estimated ± 4.8 pp, low confidence
110Nemotron 3 Nano 30B57.7%estimated ± 4.8 pp, low confidence
111Nemotron 3 Super 100B57.7%estimated ± 4.8 pp, low confidence
112Nemotron Ultra 253B57.7%estimated ± 4.8 pp, low confidence
113North Mini Code57.7%estimated ± 4.8 pp, low confidence
114Nova Pro57.7%estimated ± 4.8 pp, low confidence
115o357.7%estimated ± 4.8 pp, low confidence
116Phi-457.7%estimated ± 4.8 pp, low confidence
117Phi-4 Multimodal Instruct57.7%estimated ± 4.8 pp, low confidence
118Quasar 438B57.7%estimated ± 4.8 pp, low confidence
119Qwen2.5 Coder 32B Instruct57.7%estimated ± 4.8 pp, low confidence
120Qwen3.5 397B (Reasoning)57.7%estimated ± 4.8 pp, low confidence
121Qwen 3.6 Max (preview)57.7%estimated ± 4.8 pp, low confidence
122Qwen3 Max57.7%estimated ± 4.8 pp, low confidence
123Qwen3-Omni-30B-A3B-Instruct57.7%estimated ± 4.8 pp, low confidence
124Qwen3-Omni-30B-A3B-Thinking57.7%estimated ± 4.8 pp, low confidence
125Sarvam 105B57.7%estimated ± 4.8 pp, low confidence
126Sarvam 30B57.7%estimated ± 4.8 pp, low confidence
127Solar Pro 257.7%estimated ± 4.8 pp, low confidence
128Solar Pro 357.7%estimated ± 4.8 pp, low confidence
129Step 3.7 Flash57.7%estimated ± 4.8 pp, low confidence
130Ultravox v0.6 Llama 3.3 70B57.7%estimated ± 4.8 pp, low confidence
131Pareto 26.957.0%estimated ± 4.8 pp, medium confidence
132Step 5 Preview56.9%estimated ± 4.8 pp, medium confidence
133Claude Opus 4.7 (Adaptive)56.8%estimated ± 4.3 pp, medium confidence
134GPT-5.5 Pro56.8%estimated ± 4.6 pp, high confidence
135Grok 4.756.7%measured
136Grok 4.2056.6%estimated ± 1.3 pp, medium confidence
137Inkling56.6%estimated ± 1.3 pp, medium confidence
138GPT-5.4 Pro56.4%estimated ± 4.6 pp, high confidence
139DeepSeek V4 Flash 073156.3%estimated ± 1.3 pp, medium confidence
140Gemini 3.1 Flash-Lite56.3%estimated ± 1.3 pp, medium confidence
141GLM-5.3-Flash56.0%estimated ± 1.3 pp, medium confidence
142dots3-note Preview55.8%estimated ± 4.6 pp, high confidence
143GPT-5.6 Luna55.7%measured
144Ornith-1.5-397B55.6%estimated ± 4.8 pp, medium confidence
145Claude Opus 4.6 (Adaptive)55.1%estimated ± 4.6 pp, high confidence
146Claude Opus 4.5 Thinking55.1%estimated ± 4.6 pp, medium confidence
147Claude Sonnet 4.555.1%estimated ± 4.6 pp, medium confidence
148Claude Sonnet 4.5 Thinking55.1%estimated ± 4.6 pp, medium confidence
149Gemini 3 Pro55.1%estimated ± 4.6 pp, medium confidence
150Gemini 3 Pro Deep Think55.1%estimated ± 4.6 pp, medium confidence
151GPT-5.255.1%estimated ± 4.6 pp, medium confidence
152o1-preview55.1%estimated ± 4.9 pp, low confidence
153o3-pro55.1%estimated ± 4.9 pp, low confidence
154Gemini 3.5 Flash-Lite55.0%estimated ± 1.3 pp, low confidence
155Grok 4.355.0%estimated ± 1.3 pp, low confidence
156Nemotron 3 Ultra55.0%estimated ± 1.3 pp, low confidence
157Pareto 26.10 Preview54.8%estimated ± 4.8 pp, medium confidence
158Hy4 preview54.6%estimated ± 4.8 pp, medium confidence
159Claude Haiku 5.554.5%estimated ± 4.8 pp, medium confidence
160Inkling-Small54.4%estimated ± 1.3 pp, low confidence
161Qwen3.8-Flash-Next53.5%estimated ± 4.8 pp, low confidence
162Claude Opus 4.652.7%estimated ± 4.8 pp, low confidence
163Qwen3.8-Omni-Flash52.1%estimated ± 4.8 pp, low confidence
164DeepSeek V4.1 Flash51.9%estimated ± 4.8 pp, low confidence
165GPT-5.4 mini51.7%estimated ± 1.3 pp, low confidence
166MiMo-V2.5-Pro51.7%estimated ± 1.3 pp, low confidence
167Qwen3.8-27B51.2%estimated ± 1.3 pp, low confidence
168Beam51.1%estimated ± 4.8 pp, low confidence
169MiniMax M351.0%estimated ± 1.3 pp, low confidence
170Qwen3.5 Flash50.9%estimated ± 1.3 pp, low confidence
171Qwen3.7 Plus50.7%estimated ± 4.8 pp, low confidence
172MiMo-V2.550.1%estimated ± 1.3 pp, low confidence
173GLM-4.750.0%estimated ± 1.3 pp, low confidence
174Interfaze Beta50.0%estimated ± 4.8 pp, low confidence
175GLM-4.649.9%estimated ± 1.3 pp, low confidence
176Ling 3.0 Flash49.9%estimated ± 1.3 pp, low confidence
177GLM-4.549.9%estimated ± 1.3 pp, low confidence
178MiniMax M2.749.8%estimated ± 1.3 pp, low confidence
179Claude Haiku 4.549.8%estimated ± 1.3 pp, low confidence
180GPT-5.4 nano49.8%estimated ± 1.3 pp, low confidence
181Laguna M.149.8%estimated ± 1.3 pp, low confidence
182Laguna XS.249.8%estimated ± 1.3 pp, low confidence
183Mistral Medium 3.5 128B49.8%estimated ± 1.3 pp, low confidence
184Ornith-1.5-35B-A3B48.6%estimated ± 4.8 pp, low confidence
185Solar Pro 448.2%estimated ± 4.8 pp, low confidence
186GPT-5.448.1%measured
187Qwen3.5 397B47.0%estimated ± 4.8 pp, low confidence
188Qwen3.6-27B45.9%estimated ± 4.8 pp, low confidence
189Kimi K2.545.5%estimated ± 4.8 pp, low confidence
190Kimi K2.5 (Reasoning)45.5%estimated ± 4.8 pp, low confidence
191Hy3 Preview44.7%estimated ± 4.8 pp, low confidence
192Claude Opus 4.544.3%estimated ± 4.8 pp, low confidence
193Qwen3.5-122B-A10B43.6%estimated ± 4.8 pp, low confidence
194Ornith-1.5-9B43.2%estimated ± 4.8 pp, low confidence
195Solar Open 243.0%estimated ± 4.8 pp, low confidence
196GLM-542.4%estimated ± 4.8 pp, low confidence
197Qwen3.6-35B-A3B42.4%estimated ± 4.8 pp, low confidence
198MiMo-V2-Flash42.0%estimated ± 1.7 pp, low confidence
199Ternary Bonsai 2 27B42.0%estimated ± 4.8 pp, low confidence
200A.X K241.7%estimated ± 4.8 pp, low confidence
201Qwen3.5-27B41.5%estimated ± 4.8 pp, low confidence
202Gemma 4 31B39.2%estimated ± 4.8 pp, low confidence
203MAI-Thinking-139.0%estimated ± 4.8 pp, low confidence
204Qwen3.5-35B-A3B39.0%estimated ± 4.8 pp, low confidence
205Ling 3.0 Flash FP838.6%estimated ± 4.8 pp, low confidence
206Gemini 2.5 Pro36.8%estimated ± 4.8 pp, low confidence
207K-EXAONE 2.035.3%estimated ± 4.8 pp, low confidence
208Mercury 2.529.7%estimated ± 4.8 pp, low confidence
209o1-pro29.7%estimated ± 4.8 pp, low confidence
210Gemma 4 12B29.3%estimated ± 4.8 pp, low confidence
211Qwen3 235B 250727.2%estimated ± 4.8 pp, low confidence
212o3-mini26.7%estimated ± 4.8 pp, low confidence
213Trinity-Large-Thinking25.3%estimated ± 4.8 pp, low confidence
214o124.3%estimated ± 4.8 pp, low confidence
215Nemotron 3.5 Lightning 30B A3B NVFP424.1%estimated ± 4.8 pp, low confidence
216Nemotron 3 Nano Omni 30B A3B19.3%estimated ± 4.8 pp, low confidence
217ZAYA1-8B17.7%estimated ± 4.8 pp, low confidence
218MiniCPM5-2B16.7%estimated ± 4.8 pp, low confidence
219LongCat-Flash-Lite-Sparse15.8%estimated ± 4.8 pp, low confidence
220Gemma 4 26B A4B13.7%estimated ± 4.8 pp, low confidence
221Granite 4.2 30B12.5%estimated ± 4.8 pp, low confidence
222GPT-4.112.4%estimated ± 4.8 pp, low confidence
223GPT-4.1 mini10.4%estimated ± 4.8 pp, low confidence
224Granite 4.2 8B10.3%estimated ± 4.8 pp, low confidence
225Trinity-Large-Preview9.6%estimated ± 4.8 pp, low confidence
226Claude 3.5 Sonnet6.7%estimated ± 4.8 pp, low confidence
227DeepSeek V36.5%estimated ± 4.8 pp, low confidence
228Ling 2.6 Flash6.5%estimated ± 4.8 pp, low confidence
229Gemma 4 E4B6.2%estimated ± 4.8 pp, low confidence
230Mellum2-12B-A2.5B-Thinking5.6%estimated ± 4.8 pp, low confidence
231ZAYA1-74B-Preview5.5%estimated ± 4.8 pp, low confidence
232Granite 4.2 3B4.2%estimated ± 4.8 pp, low confidence
233GPT-4.1 nano2.6%estimated ± 4.8 pp, low confidence
234LLaDA2.2-mini1.2%estimated ± 4.8 pp, low confidence
235Gemma 4 E2B1.1%estimated ± 4.8 pp, low confidence
236Soofi S 30B-A3B1.1%estimated ± 4.8 pp, low confidence
237Mellum2-12B-A2.5B-Instruct0.8%estimated ± 4.8 pp, low confidence
238MiniCPM5-1B0.1%estimated ± 4.8 pp, low confidence
239LFM2.5-VL-450M0.0%estimated ± 4.8 pp, low confidence
240LFM2.5-230M0.0%estimated ± 4.8 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General