benchgap
Knowledge & reasoning

HealthBench Hard leaderboard

As of 2026-10-07, the highest measured score on HealthBench Hard is 42.8% by Muse Spark. 213 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Muse Spark42.8%measured
2GPT-5.440.1%measured
3Sakana Fugu36.6%estimated ± 5.1 pp, low confidence
4Sakana Fugu-Ultra36.6%estimated ± 5.1 pp, low confidence
5GPT-6 Astra36.6%measured
6Claude Opus 4.7 (Adaptive)36.5%estimated ± 5.1 pp, low confidence
7Claude Mythos 536.4%estimated ± 5.1 pp, low confidence
8GPT-6.1 Sol36.2%measured
9Claude Opus 4.836.1%estimated ± 5.1 pp, low confidence
10GPT-5.536.1%estimated ± 5.1 pp, low confidence
11Kimi K336.0%estimated ± 5.1 pp, low confidence
12Claude Opus 5.535.2%estimated ± 7.7 pp, low confidence
13Claude Sonnet 5.535.1%estimated ± 7.7 pp, low confidence
14Claude Fable 5.135.0%estimated ± 7.7 pp, low confidence
15Gemini 4 Argon35.0%estimated ± 7.7 pp, low confidence
16Claude Opus 534.9%estimated ± 7.7 pp, low confidence
17Claude Fable 534.8%estimated ± 7.7 pp, low confidence
18Muse Spark 1.334.7%estimated ± 7.7 pp, low confidence
19Grok 4.734.6%estimated ± 7.7 pp, low confidence
20MiMo-V2.6-Pro34.6%estimated ± 7.7 pp, low confidence
21Qwen3.8 Max Preview34.5%estimated ± 7.7 pp, low confidence
22Ornith-1.5-397B34.4%estimated ± 5.1 pp, low confidence
23GLM-5.334.4%estimated ± 7.7 pp, low confidence
24Grok 4.634.3%estimated ± 7.7 pp, low confidence
25GPT-5.5 Pro34.3%estimated ± 9.0 pp, low confidence
26Step 5 Preview34.2%estimated ± 7.7 pp, low confidence
27GPT-5.4 Pro34.2%estimated ± 9.0 pp, low confidence
28Claude Haiku 5.534.2%estimated ± 7.7 pp, low confidence
29GLM-5.3-Flash33.9%estimated ± 7.7 pp, low confidence
30Ling 3.1 Flash33.7%estimated ± 7.7 pp, low confidence
31Gemini 3 Pro Deep Think33.7%estimated ± 9.0 pp, low confidence
32Gemini 3.8 Flash33.7%estimated ± 7.7 pp, low confidence
33Qwen3.8 Max33.4%estimated ± 5.1 pp, low confidence
34Muse Spark 1.233.3%estimated ± 7.7 pp, low confidence
35Gemini 3.7 Flash33.2%estimated ± 7.7 pp, low confidence
36GPT-5.6 Sol33.1%measured
37Grok 4.533.1%estimated ± 7.7 pp, low confidence
38Mistral Large 432.9%estimated ± 7.7 pp, low confidence
39Claude Sonnet 532.9%estimated ± 7.7 pp, low confidence
40MiMo-V2.6-Flash32.8%estimated ± 7.7 pp, low confidence
41GPT-5.6 Terra32.7%measured
42GPT-5.232.1%estimated ± 5.1 pp, low confidence
43Qwen3.7 Max32.1%estimated ± 5.1 pp, low confidence
44GPT-5.6 Luna32.0%measured
45GPT-6 Luna31.4%measured
46Hy4 preview31.2%estimated ± 5.1 pp, low confidence
47Gemini 3.6 Flash30.7%estimated ± 7.7 pp, low confidence
48Muse Spark 1.130.5%estimated ± 7.7 pp, low confidence
49Gemini 3.5 Flash30.2%estimated ± 5.1 pp, low confidence
50GPT-6 Sol30.1%measured
51GPT-5.3 Codex29.5%estimated ± 7.7 pp, low confidence
52Claude Opus 4.6 (Adaptive)28.9%estimated ± 7.7 pp, low confidence
53Claude Opus 4.727.8%estimated ± 7.7 pp, low confidence
54MiniMax M325.6%estimated ± 7.7 pp, low confidence
55Claude Opus 4.5 Thinking25.5%estimated ± 7.7 pp, low confidence
56MiMo-V2-Pro24.8%estimated ± 7.7 pp, low confidence
57GPT-5.2-Codex24.5%estimated ± 7.7 pp, low confidence
58Qwen 3.6 Max (preview)24.3%estimated ± 7.7 pp, low confidence
59Solar Pro 424.0%estimated ± 7.7 pp, low confidence
60Gemini 3 Pro23.7%estimated ± 7.7 pp, low confidence
61Qwen3.8-Flash-Next23.3%estimated ± 5.1 pp, low confidence
62Quasar 438B21.5%estimated ± 7.7 pp, low confidence
63GLM-5-Turbo21.2%estimated ± 7.7 pp, low confidence
64Apodex 1.120.8%estimated ± 7.7 pp, low confidence
65Apodex 1.1 Mini20.8%estimated ± 7.7 pp, low confidence
66Gemini 3.1 Pro20.6%measured
67Grok 4.2020.3%measured
68GLM-5.120.1%estimated ± 7.7 pp, low confidence
69MiMo-V2.5-Pro20.0%estimated ± 7.7 pp, low confidence
70Kimi K2.7 Code19.6%estimated ± 7.7 pp, low confidence
71Hy318.6%estimated ± 7.7 pp, low confidence
72GPT-5.117.4%estimated ± 7.7 pp, low confidence
73Ling 3.0 Flash VL17.0%estimated ± 7.7 pp, low confidence
74MiMo-V2-Omni15.6%estimated ± 7.7 pp, low confidence
75GPT-5.1-Codex15.1%estimated ± 7.7 pp, low confidence
76GPT-5.1-Codex-Max15.1%estimated ± 7.7 pp, low confidence
77Claude Opus 4.614.8%measured
78GLM-5V-Turbo14.7%estimated ± 7.7 pp, low confidence
79GLM-5.214.3%estimated ± 5.1 pp, low confidence
80GPT-5 (high)13.6%estimated ± 7.7 pp, low confidence
81GPT-5 (medium)13.3%estimated ± 7.7 pp, low confidence
82Claude 4.1 Opus Thinking13.3%estimated ± 7.7 pp, low confidence
83MiniMax M2.713.1%estimated ± 7.7 pp, low confidence
84Command A+12.5%estimated ± 7.7 pp, low confidence
85Grok 412.5%estimated ± 7.7 pp, low confidence
86Gemini 3.5 Flash-Lite11.8%estimated ± 7.7 pp, low confidence
87o3-pro11.2%estimated ± 7.7 pp, low confidence
88Qwen3.8-Omni-Flash11.0%estimated ± 5.1 pp, low confidence
89A.X K210.3%estimated ± 7.7 pp, low confidence
90Qwen3.5 397B (Reasoning)10.3%estimated ± 7.7 pp, low confidence
91DeepSeek V4.1 Flash9.5%estimated ± 5.1 pp, low confidence
92Grok 4.1 Fast (Reasoning)8.2%estimated ± 7.7 pp, low confidence
93o37.9%estimated ± 7.7 pp, low confidence
94K-EXAONE 2.07.0%estimated ± 7.7 pp, low confidence
95Step 3.7 Flash6.6%estimated ± 7.7 pp, low confidence
96Claude 4.1 Opus5.2%estimated ± 7.7 pp, low confidence
97Kimi K2.65.0%estimated ± 5.1 pp, low confidence
98Grok 4 Fast (Reasoning)4.4%estimated ± 7.7 pp, low confidence
99Gemini 3 Flash4.3%estimated ± 7.7 pp, low confidence
100Qwen3.6 Plus4.2%estimated ± 5.1 pp, low confidence
101Muse Glimmer 30B3.8%estimated ± 7.7 pp, low confidence
102Qwen3.7 Plus3.5%estimated ± 5.1 pp, low confidence
103Gemma 4 26B A4B2.9%estimated ± 7.7 pp, low confidence
104Claude 4 Sonnet2.9%estimated ± 7.7 pp, low confidence
105DeepSeek V4 Pro 08132.4%estimated ± 5.1 pp, low confidence
106Grok 4.32.4%estimated ± 5.1 pp, low confidence
107DeepSeek V3.22.4%estimated ± 7.7 pp, low confidence
108Qwen3 Max2.0%estimated ± 7.7 pp, low confidence
109Claude Sonnet 4.61.7%estimated ± 5.1 pp, low confidence
110Interfaze Beta1.7%estimated ± 5.1 pp, low confidence
111GLM-4.61.6%estimated ± 7.7 pp, low confidence
112K-Exaone1.3%estimated ± 7.7 pp, low confidence
113Mistral Medium 3.5 128B1.2%estimated ± 7.7 pp, low confidence
114Grok Code Fast 11.1%estimated ± 7.7 pp, low confidence
115DeepSeek V3.11.0%estimated ± 7.7 pp, low confidence
116DeepSeek V3.1 (Reasoning)0.9%estimated ± 7.7 pp, low confidence
117Inkling-Small0.8%estimated ± 5.1 pp, low confidence
118DeepSeek-R10.7%estimated ± 7.7 pp, low confidence
119Nemotron 3 Super 100B0.7%estimated ± 7.7 pp, low confidence
120Kimi K20.6%estimated ± 7.7 pp, low confidence
121MiniCPM5-2B0.5%estimated ± 7.7 pp, low confidence
122Mercury 2.50.5%estimated ± 7.7 pp, low confidence
123Ornith-1.5-35B-A3B0.4%estimated ± 5.1 pp, low confidence
124Qwen3.8-27B0.4%estimated ± 5.1 pp, low confidence
125GPT-OSS 120B0.4%estimated ± 7.7 pp, low confidence
126o1-preview0.3%estimated ± 7.7 pp, low confidence
127Grok 4.1 Fast0.3%estimated ± 7.7 pp, low confidence
128Mistral Small 40.3%estimated ± 7.7 pp, low confidence
129Mistral Small 4 (Reasoning)0.3%estimated ± 7.7 pp, low confidence
130GLM-4.5-Air0.3%estimated ± 7.7 pp, low confidence
131Ling 3.0 Tiny0.3%estimated ± 7.7 pp, low confidence
132Trinity-Large-Preview0.2%estimated ± 7.7 pp, low confidence
133Trinity-Large-Thinking0.2%estimated ± 7.7 pp, low confidence
134Llama 4 Maverick0.1%estimated ± 7.7 pp, low confidence
135North Mini Code0.1%estimated ± 7.7 pp, low confidence
136Gemini 2.5 Flash0.1%estimated ± 7.7 pp, low confidence
137DeepSeek V3 03240.1%estimated ± 7.7 pp, low confidence
138Mistral Large 30.1%estimated ± 7.7 pp, low confidence
139Qwen3.5 397B0.1%estimated ± 5.1 pp, low confidence
140Mistral Medium 30.1%estimated ± 7.7 pp, low confidence
141GPT-OSS 20B0.1%estimated ± 7.7 pp, low confidence
142Nemotron 3 Nano 30B0.1%estimated ± 7.7 pp, low confidence
143Sarvam 105B0.1%estimated ± 7.7 pp, low confidence
144Claude 3 Opus0.1%estimated ± 7.7 pp, low confidence
145GPT-4o0.1%estimated ± 7.7 pp, low confidence
146LFM2.5-2.6B0.1%estimated ± 7.7 pp, low confidence
147DeepSeek R1 Distill Qwen 32B0.1%estimated ± 7.7 pp, low confidence
148DeepSeek V4 Flash 07310.0%estimated ± 5.1 pp, low confidence
149Llama 4 Scout0.0%estimated ± 7.7 pp, low confidence
150GPT-5.4 mini0.0%estimated ± 5.1 pp, low confidence
151Gemini 1.5 Pro0.0%estimated ± 7.7 pp, low confidence
152Solar Pro 30.0%estimated ± 7.7 pp, low confidence
153Qwen3-Omni-30B-A3B-Thinking0.0%estimated ± 7.7 pp, low confidence
154Inkling0.0%estimated ± 5.1 pp, low confidence
155Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 7.7 pp, low confidence
156Mistral Large 20.0%estimated ± 7.7 pp, low confidence
157Nemotron Ultra 253B0.0%estimated ± 7.7 pp, low confidence
158Qwen3.6-27B0.0%estimated ± 5.1 pp, low confidence
159Llama 3.1 405B0.0%estimated ± 7.7 pp, low confidence
160LFM2.5-8B-A1B0.0%estimated ± 7.7 pp, low confidence
161GPT-4 Turbo0.0%estimated ± 7.7 pp, low confidence
162Kimi K2.50.0%estimated ± 5.1 pp, low confidence
163Kimi K2.5 (Reasoning)0.0%estimated ± 5.1 pp, low confidence
164Solar Pro 20.0%estimated ± 7.7 pp, low confidence
165Nova Pro0.0%estimated ± 7.7 pp, low confidence
166Qwen2.5 Coder 32B Instruct0.0%estimated ± 7.7 pp, low confidence
167GPT-4o mini0.0%estimated ± 7.7 pp, low confidence
168Sarvam 30B0.0%estimated ± 7.7 pp, low confidence
169Celeris-10.0%estimated ± 7.7 pp, low confidence
170Exaone 4.0 32B0.0%estimated ± 7.7 pp, low confidence
171Hy3 Preview0.0%estimated ± 5.1 pp, low confidence
172Qwen3-Omni-30B-A3B-Instruct0.0%estimated ± 7.7 pp, low confidence
173Phi-40.0%estimated ± 7.7 pp, low confidence
174Phi-4 Multimodal Instruct0.0%estimated ± 7.7 pp, low confidence
175Claude Opus 4.50.0%estimated ± 5.1 pp, low confidence
176Nemotron 3 Ultra0.0%estimated ± 5.1 pp, low confidence
177Claude 3 Haiku0.0%estimated ± 7.7 pp, low confidence
178Gemini 1.0 Pro0.0%estimated ± 7.7 pp, low confidence
179Exaone 4.0 1.2B0.0%estimated ± 7.7 pp, low confidence
180Granite-4.0-H-1B0.0%estimated ± 7.7 pp, low confidence
181Qwen3.5-122B-A10B0.0%estimated ± 5.1 pp, low confidence
182Gemma 3 27B0.0%estimated ± 7.7 pp, low confidence
183Granite-4.0-350M0.0%estimated ± 7.7 pp, low confidence
184Granite-4.0-H-350M0.0%estimated ± 7.7 pp, low confidence
185LFM2.5-VL-1.6B-Extract0.0%estimated ± 7.7 pp, low confidence
186Ornith-1.5-9B0.0%estimated ± 5.1 pp, low confidence
187GLM-50.0%estimated ± 5.1 pp, low confidence
188Qwen3.6-35B-A3B0.0%estimated ± 5.1 pp, low confidence
189GLM-4.70.0%estimated ± 5.1 pp, low confidence
190Ternary Bonsai 2 27B0.0%estimated ± 5.1 pp, low confidence
191Qwen3.5-27B0.0%estimated ± 5.1 pp, low confidence
192Ling 3.0 Flash0.0%estimated ± 5.1 pp, low confidence
193Claude 3.5 Sonnet0.0%estimated ± 5.1 pp, low confidence
194Claude Sonnet 4.50.0%estimated ± 5.1 pp, low confidence
195DeepSeek V30.0%estimated ± 5.1 pp, low confidence
196Gemini 2.5 Pro0.0%estimated ± 5.1 pp, low confidence
197Gemma 4 12B0.0%estimated ± 5.1 pp, low confidence
198Gemma 4 31B0.0%estimated ± 5.1 pp, low confidence
199Gemma 4 E2B0.0%estimated ± 5.1 pp, low confidence
200Gemma 4 E4B0.0%estimated ± 5.1 pp, low confidence
201GPT-4.10.0%estimated ± 5.1 pp, low confidence
202GPT-4.1 mini0.0%estimated ± 5.1 pp, low confidence
203GPT-4.1 nano0.0%estimated ± 5.1 pp, low confidence
204GPT-5.4 nano0.0%estimated ± 5.1 pp, low confidence
205Granite 4.2 30B0.0%estimated ± 5.1 pp, low confidence
206Granite 4.2 3B0.0%estimated ± 5.1 pp, low confidence
207Granite 4.2 8B0.0%estimated ± 5.1 pp, low confidence
208LFM2.5-230M0.0%estimated ± 5.1 pp, low confidence
209LFM2.5-VL-450M0.0%estimated ± 5.1 pp, low confidence
210Ling 2.6 Flash0.0%estimated ± 5.1 pp, low confidence
211Ling 3.0 Flash FP80.0%estimated ± 5.1 pp, low confidence
212MAI-Thinking-10.0%estimated ± 5.1 pp, low confidence
213Mellum2-12B-A2.5B-Instruct0.0%estimated ± 5.1 pp, low confidence
214Mellum2-12B-A2.5B-Thinking0.0%estimated ± 5.1 pp, low confidence
215MiMo-V2-Flash0.0%estimated ± 5.1 pp, low confidence
216Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 5.1 pp, low confidence
217Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 5.1 pp, low confidence
218o10.0%estimated ± 5.1 pp, low confidence
219o1-pro0.0%estimated ± 5.1 pp, low confidence
220o3-mini0.0%estimated ± 5.1 pp, low confidence
221Qwen3 235B 25070.0%estimated ± 5.1 pp, low confidence
222Qwen3.5-35B-A3B0.0%estimated ± 5.1 pp, low confidence
223Soofi S 30B-A3B0.0%estimated ± 5.1 pp, low confidence
224ZAYA1-74B-Preview0.0%estimated ± 5.1 pp, low confidence
225ZAYA1-8B0.0%estimated ± 5.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General