benchgap
Knowledge & reasoning

HealthBench Professional (raw) leaderboard

As of 2026-10-07, the highest measured score on HealthBench Professional (raw) is 77.1% by Claude Opus 5.5. 193 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.577.1%measured
2Claude Sonnet 5.577.1%measured
3Claude Fable 5.174.3%estimated ± 5.0 pp, medium confidence
4Claude Opus 573.4%measured
5Gemini 4 Argon72.9%estimated ± 5.0 pp, medium confidence
6Claude Fable 571.7%estimated ± 5.0 pp, medium confidence
7GPT-6 Astra68.2%measured
8GPT-5.6 Sol67.2%estimated ± 5.0 pp, medium confidence
9GPT-6.1 Sol67.2%measured
10MiMo-V2.6-Pro67.1%estimated ± 5.0 pp, medium confidence
11Claude Opus 4.866.6%estimated ± 5.0 pp, medium confidence
12Muse Spark 1.366.6%estimated ± 5.0 pp, medium confidence
13Gemini 3.7 Flash66.0%estimated ± 5.0 pp, medium confidence
14Gemini 3.8 Flash65.9%estimated ± 5.0 pp, medium confidence
15Gemini 3.1 Pro65.2%estimated ± 5.0 pp, medium confidence
16Kimi K365.2%estimated ± 5.0 pp, medium confidence
17Step 5 Preview64.8%estimated ± 5.0 pp, medium confidence
18Muse Spark 1.164.6%estimated ± 5.0 pp, medium confidence
19GPT-5.564.3%estimated ± 5.0 pp, medium confidence
20Muse Spark 1.264.0%estimated ± 5.0 pp, medium confidence
21Claude Haiku 5.563.1%estimated ± 5.0 pp, medium confidence
22GPT-5.462.6%estimated ± 5.0 pp, medium confidence
23Grok 4.762.1%estimated ± 5.0 pp, medium confidence
24Qwen3.8 Max Preview62.1%estimated ± 5.0 pp, medium confidence
25Grok 4.2062.0%estimated ± 5.3 pp, low confidence
26GPT-5.6 Terra61.9%estimated ± 5.0 pp, medium confidence
27Grok 4.661.9%estimated ± 5.0 pp, medium confidence
28Gemini 3.5 Flash61.7%estimated ± 5.0 pp, medium confidence
29Grok 4.561.7%estimated ± 5.0 pp, medium confidence
30GPT-5.3 Codex61.6%estimated ± 5.0 pp, medium confidence
31Claude Opus 4.7 (Adaptive)61.4%estimated ± 5.0 pp, medium confidence
32GLM-5.361.4%estimated ± 5.0 pp, medium confidence
33GPT-6 Luna61.2%measured
34Claude Sonnet 560.5%estimated ± 5.0 pp, medium confidence
35GLM-5.260.4%estimated ± 5.0 pp, medium confidence
36DeepSeek V4 Pro 081360.3%estimated ± 5.0 pp, medium confidence
37Gemini 3.6 Flash60.1%estimated ± 5.0 pp, medium confidence
38Muse Spark60.0%estimated ± 5.0 pp, medium confidence
39Qwen3.7 Max59.9%estimated ± 5.0 pp, medium confidence
40GPT-6 Sol59.5%measured
41Claude Opus 4.6 (Adaptive)59.3%estimated ± 5.0 pp, medium confidence
42GLM-5.3-Flash59.3%estimated ± 5.0 pp, medium confidence
43Gemini 3 Pro59.2%estimated ± 5.0 pp, medium confidence
44GPT-5.6 Luna59.0%estimated ± 5.0 pp, medium confidence
45Ling 3.1 Flash58.9%estimated ± 5.0 pp, medium confidence
46DeepSeek V4.1 Flash58.7%estimated ± 5.0 pp, medium confidence
47MiniMax M358.6%estimated ± 5.0 pp, medium confidence
48DeepSeek V4 Flash 073158.2%estimated ± 5.0 pp, medium confidence
49Qwen3.8-Flash-Next57.7%estimated ± 5.0 pp, low confidence
50GPT-5.257.4%estimated ± 5.0 pp, low confidence
51Kimi K2.657.2%estimated ± 5.0 pp, low confidence
52Grok 4.357.0%estimated ± 5.0 pp, low confidence
53GPT-5.2-Codex55.6%estimated ± 5.0 pp, low confidence
54MiMo-V2.5-Pro55.6%estimated ± 5.0 pp, low confidence
55Qwen3.7 Plus55.5%estimated ± 5.0 pp, low confidence
56MiMo-V2.6-Flash55.1%estimated ± 5.0 pp, low confidence
57Kimi K2.7 Code55.0%estimated ± 5.0 pp, low confidence
58Mistral Large 455.0%estimated ± 5.0 pp, low confidence
59Apodex 1.154.2%estimated ± 5.0 pp, low confidence
60Apodex 1.1 Mini54.2%estimated ± 5.0 pp, low confidence
61Qwen3.8-27B54.0%estimated ± 5.0 pp, low confidence
62Hy353.6%estimated ± 5.0 pp, low confidence
63Hy3 Preview53.6%estimated ± 5.0 pp, low confidence
64Claude Opus 4.753.4%estimated ± 5.0 pp, low confidence
65Inkling-Small53.4%estimated ± 5.0 pp, low confidence
66Inkling52.1%estimated ± 5.0 pp, low confidence
67Qwen 3.6 Max (preview)51.1%estimated ± 5.0 pp, low confidence
68Kimi K2.551.0%estimated ± 5.0 pp, low confidence
69Kimi K2.5 (Reasoning)51.0%estimated ± 5.0 pp, low confidence
70MiMo-V2-Pro50.7%estimated ± 5.0 pp, low confidence
71Claude Opus 4.5 Thinking50.4%estimated ± 5.0 pp, low confidence
72GLM-5.150.4%estimated ± 5.0 pp, low confidence
73A.X K249.9%estimated ± 5.0 pp, low confidence
74MiniMax M2.749.9%estimated ± 5.0 pp, low confidence
75GLM-549.6%estimated ± 5.0 pp, low confidence
76Solar Pro 449.5%estimated ± 5.0 pp, low confidence
77GPT-5.148.8%estimated ± 5.0 pp, low confidence
78GPT-5 (high)48.8%estimated ± 5.0 pp, low confidence
79Nemotron 3 Ultra48.7%estimated ± 5.0 pp, low confidence
80GPT-5.4 nano48.6%estimated ± 5.0 pp, low confidence
81GPT-5.4 mini48.4%estimated ± 5.0 pp, low confidence
82GLM-5-Turbo48.1%estimated ± 5.0 pp, low confidence
83Qwen3.6 Plus48.1%estimated ± 5.0 pp, low confidence
84GLM-4.747.7%estimated ± 5.0 pp, low confidence
85Grok 447.0%estimated ± 5.0 pp, low confidence
86GPT-5.1-Codex46.0%estimated ± 5.0 pp, low confidence
87GPT-5.1-Codex-Max46.0%estimated ± 5.0 pp, low confidence
88GPT-5 (medium)45.7%estimated ± 5.0 pp, low confidence
89Qwen3.5-122B-A10B45.5%estimated ± 5.0 pp, low confidence
90Qwen3.5-27B44.2%estimated ± 5.0 pp, low confidence
91Ling 3.0 Flash44.0%estimated ± 5.0 pp, low confidence
92Ling 3.0 Flash FP844.0%estimated ± 5.0 pp, low confidence
93Gemma 4 31B43.9%estimated ± 5.0 pp, low confidence
94Qwen3.6-27B43.3%estimated ± 5.0 pp, low confidence
95Gemini 2.5 Pro42.7%estimated ± 5.0 pp, low confidence
96Qwen3.6-35B-A3B42.4%estimated ± 5.0 pp, low confidence
97MiMo-V2-Omni42.3%estimated ± 5.0 pp, low confidence
98Ling 3.0 Flash VL42.2%estimated ± 5.0 pp, low confidence
99Muse Glimmer 30B42.2%estimated ± 5.0 pp, low confidence
100Step 3.7 Flash41.5%estimated ± 5.0 pp, low confidence
101Qwen3.5-35B-A3B41.1%estimated ± 5.0 pp, low confidence
102Nemotron 3 Super 100B40.9%estimated ± 5.0 pp, low confidence
103o3-pro40.8%estimated ± 5.0 pp, low confidence
104o340.1%estimated ± 5.0 pp, low confidence
105Qwen3.5 397B39.8%estimated ± 5.0 pp, low confidence
106Qwen3.5 397B (Reasoning)39.8%estimated ± 5.0 pp, low confidence
107GPT-OSS 120B39.6%estimated ± 5.0 pp, low confidence
108Gemma 4 26B A4B39.3%estimated ± 5.0 pp, low confidence
109Grok 4.1 Fast (Reasoning)39.3%estimated ± 5.0 pp, low confidence
110Claude Opus 4.639.0%estimated ± 5.0 pp, low confidence
111Grok 4 Fast (Reasoning)39.0%estimated ± 5.0 pp, low confidence
112Gemini 3.5 Flash-Lite38.7%estimated ± 5.0 pp, low confidence
113Quasar 438B38.6%estimated ± 5.0 pp, low confidence
114K-EXAONE 2.038.5%estimated ± 5.0 pp, low confidence
115Claude 4.1 Opus36.8%estimated ± 5.0 pp, low confidence
116GLM-5V-Turbo36.8%estimated ± 5.0 pp, low confidence
117DeepSeek-R135.3%estimated ± 5.0 pp, low confidence
118Trinity-Large-Preview35.3%estimated ± 5.0 pp, low confidence
119Trinity-Large-Thinking35.3%estimated ± 5.0 pp, low confidence
120Gemma 4 12B35.2%estimated ± 5.0 pp, low confidence
121Gemini 3 Flash34.4%estimated ± 5.0 pp, low confidence
122DeepSeek V3.1 (Reasoning)33.6%estimated ± 5.0 pp, low confidence
123K-Exaone33.1%estimated ± 5.0 pp, low confidence
124Mistral Medium 3.5 128B33.0%estimated ± 5.0 pp, low confidence
125Claude Sonnet 4.632.4%estimated ± 5.0 pp, low confidence
126Claude Opus 4.532.3%estimated ± 5.0 pp, low confidence
127Claude 4.1 Opus Thinking31.4%estimated ± 5.0 pp, low confidence
128Command A+30.8%estimated ± 5.0 pp, low confidence
129Qwen3 Max30.7%estimated ± 5.0 pp, low confidence
130Mercury 2.530.6%estimated ± 5.0 pp, low confidence
131Nemotron 3 Nano 30B30.1%estimated ± 5.0 pp, low confidence
132DeepSeek V3.229.9%estimated ± 5.0 pp, low confidence
133Granite 4.2 30B29.9%estimated ± 5.0 pp, low confidence
134North Mini Code29.7%estimated ± 5.0 pp, low confidence
135GPT-OSS 20B29.6%estimated ± 5.0 pp, low confidence
136Sarvam 105B29.6%estimated ± 5.0 pp, low confidence
137Nemotron 3.5 Lightning 30B A3B NVFP429.1%estimated ± 5.0 pp, low confidence
138o1-pro28.8%estimated ± 5.0 pp, low confidence
139Solar Pro 328.8%estimated ± 5.0 pp, low confidence
140Mistral Small 428.3%estimated ± 5.0 pp, low confidence
141Mistral Small 4 (Reasoning)28.3%estimated ± 5.0 pp, low confidence
142Granite 4.2 8B28.0%estimated ± 5.0 pp, low confidence
143Ling 3.0 Tiny27.5%estimated ± 5.0 pp, low confidence
144o1-preview27.4%estimated ± 5.0 pp, low confidence
145MiniCPM5-2B27.0%estimated ± 5.0 pp, low confidence
146MiMo-V2-Flash26.6%estimated ± 5.0 pp, low confidence
147Grok Code Fast 125.9%estimated ± 5.0 pp, low confidence
148o3-mini25.7%estimated ± 5.0 pp, low confidence
149Qwen3-Omni-30B-A3B-Thinking25.2%estimated ± 5.0 pp, low confidence
150Sarvam 30B25.2%estimated ± 5.0 pp, low confidence
151Kimi K225.1%estimated ± 5.0 pp, low confidence
152Nemotron Ultra 253B25.1%estimated ± 5.0 pp, low confidence
153GLM-4.5-Air24.6%estimated ± 5.0 pp, low confidence
154o124.6%estimated ± 5.0 pp, low confidence
155LFM2.5-8B-A1B24.5%estimated ± 5.0 pp, low confidence
156Celeris-124.3%estimated ± 5.0 pp, low confidence
157DeepSeek V3.124.2%estimated ± 5.0 pp, low confidence
158Granite 4.2 3B24.1%estimated ± 5.0 pp, low confidence
159Granite-4.0-H-350M23.8%estimated ± 5.0 pp, low confidence
160Ling 2.6 Flash23.7%estimated ± 5.0 pp, low confidence
161LFM2.5-2.6B23.6%estimated ± 5.0 pp, low confidence
162Exaone 4.0 1.2B22.9%estimated ± 5.0 pp, low confidence
163GLM-4.622.6%estimated ± 5.0 pp, low confidence
164Granite-4.0-350M22.6%estimated ± 5.0 pp, low confidence
165Grok 4.1 Fast22.1%estimated ± 5.0 pp, low confidence
166LFM2.5-VL-1.6B-Extract22.1%estimated ± 5.0 pp, low confidence
167Exaone 4.0 32B22.0%estimated ± 5.0 pp, low confidence
168GPT-4.1 mini22.0%estimated ± 5.0 pp, low confidence
169Granite-4.0-H-1B22.0%estimated ± 5.0 pp, low confidence
170Phi-4 Multimodal Instruct22.0%estimated ± 5.0 pp, low confidence
171Llama 4 Maverick21.8%estimated ± 5.0 pp, low confidence
172Gemma 4 E2B21.7%estimated ± 5.0 pp, low confidence
173Nemotron 3 Nano Omni 30B A3B21.7%estimated ± 5.0 pp, low confidence
174DeepSeek V3 032421.6%estimated ± 5.0 pp, low confidence
175Gemini 2.5 Flash21.6%estimated ± 5.0 pp, low confidence
176DeepSeek R1 Distill Qwen 32B21.4%estimated ± 5.0 pp, low confidence
177Gemini 1.5 Pro21.4%estimated ± 5.0 pp, low confidence
178Qwen3-Omni-30B-A3B-Instruct21.4%estimated ± 5.0 pp, low confidence
179Gemma 3 27B21.2%estimated ± 5.0 pp, low confidence
180Claude 4 Sonnet21.0%estimated ± 5.0 pp, low confidence
181Gemini 1.0 Pro20.9%estimated ± 5.0 pp, low confidence
182GPT-4.120.9%estimated ± 5.0 pp, low confidence
183GPT-4o mini20.9%estimated ± 5.0 pp, low confidence
184Mistral Large 320.9%estimated ± 5.0 pp, low confidence
185Claude 3 Haiku20.8%estimated ± 5.0 pp, low confidence
186Mistral Medium 320.8%estimated ± 5.0 pp, low confidence
187Llama 3.1 405B20.6%estimated ± 5.0 pp, low confidence
188Gemma 4 E4B20.4%estimated ± 5.0 pp, low confidence
189GPT-4.1 nano20.4%estimated ± 5.0 pp, low confidence
190Llama 4 Scout20.4%estimated ± 5.0 pp, low confidence
191Phi-420.4%estimated ± 5.0 pp, low confidence
192Solar Pro 220.2%estimated ± 5.0 pp, low confidence
193Ultravox v0.6 Llama 3.3 70B20.1%estimated ± 5.0 pp, low confidence
194Qwen2.5 Coder 32B Instruct20.0%estimated ± 5.0 pp, low confidence
195Mistral Large 219.7%estimated ± 5.0 pp, low confidence
196Nova Pro19.6%estimated ± 5.0 pp, low confidence
197GPT-4 Turbo19.4%estimated ± 5.0 pp, low confidence
198DeepSeek V319.1%estimated ± 5.0 pp, low confidence
199Claude 3 Opus19.0%estimated ± 5.0 pp, low confidence
200GPT-4o18.5%estimated ± 5.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General