benchgap
Knowledge & reasoning

HealthBench (raw) leaderboard

As of 2026-10-07, the highest measured score on HealthBench (raw) is 69.4% by Claude Sonnet 5.5. 192 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.569.4%measured
2Claude Opus 5.568.1%measured
3Claude Opus 567.1%measured
4Claude Fable 5.164.9%estimated ± 6.4 pp, low confidence
5Gemini 4 Argon63.4%estimated ± 6.4 pp, low confidence
6Claude Fable 562.2%estimated ± 6.4 pp, low confidence
7GPT-5.6 Sol57.4%estimated ± 6.4 pp, low confidence
8MiMo-V2.6-Pro57.3%estimated ± 6.4 pp, low confidence
9GPT-6 Astra56.9%measured
10Claude Opus 4.856.7%estimated ± 6.4 pp, low confidence
11Muse Spark 1.356.7%estimated ± 6.4 pp, low confidence
12GPT-6.1 Sol56.7%measured
13Gemini 3.7 Flash56.1%estimated ± 6.4 pp, low confidence
14Gemini 3.8 Flash56.0%estimated ± 6.4 pp, low confidence
15Gemini 3.1 Pro55.3%estimated ± 6.4 pp, low confidence
16Kimi K355.2%estimated ± 6.4 pp, low confidence
17Step 5 Preview54.9%estimated ± 6.4 pp, low confidence
18Muse Spark 1.154.6%estimated ± 6.4 pp, low confidence
19GPT-5.554.3%estimated ± 6.4 pp, low confidence
20Muse Spark 1.254.0%estimated ± 6.4 pp, low confidence
21Claude Haiku 5.553.0%estimated ± 6.4 pp, low confidence
22GPT-5.452.4%estimated ± 6.4 pp, low confidence
23Grok 4.751.9%estimated ± 6.4 pp, low confidence
24Qwen3.8 Max Preview51.9%estimated ± 6.4 pp, low confidence
25GPT-5.6 Terra51.7%estimated ± 6.4 pp, low confidence
26Grok 4.651.7%estimated ± 6.4 pp, low confidence
27Gemini 3.5 Flash51.5%estimated ± 6.4 pp, low confidence
28Grok 4.551.5%estimated ± 6.4 pp, low confidence
29GPT-5.3 Codex51.4%estimated ± 6.4 pp, low confidence
30Claude Opus 4.7 (Adaptive)51.2%estimated ± 6.4 pp, low confidence
31GLM-5.351.2%estimated ± 6.4 pp, low confidence
32Claude Sonnet 550.3%estimated ± 6.4 pp, low confidence
33GLM-5.250.1%estimated ± 6.4 pp, low confidence
34DeepSeek V4 Pro 081350.0%estimated ± 6.4 pp, low confidence
35GPT-6 Luna50.0%measured
36Gemini 3.6 Flash49.8%estimated ± 6.4 pp, low confidence
37Muse Spark49.7%estimated ± 6.4 pp, low confidence
38Qwen3.7 Max49.5%estimated ± 6.4 pp, low confidence
39Claude Opus 4.6 (Adaptive)49.0%estimated ± 6.4 pp, low confidence
40GLM-5.3-Flash49.0%estimated ± 6.4 pp, low confidence
41Gemini 3 Pro48.8%estimated ± 6.4 pp, low confidence
42GPT-5.6 Luna48.6%estimated ± 6.4 pp, low confidence
43Ling 3.1 Flash48.5%estimated ± 6.4 pp, low confidence
44DeepSeek V4.1 Flash48.3%estimated ± 6.4 pp, low confidence
45MiniMax M348.1%estimated ± 6.4 pp, low confidence
46DeepSeek V4 Flash 073147.8%estimated ± 6.4 pp, low confidence
47Qwen3.8-Flash-Next47.2%estimated ± 6.4 pp, low confidence
48GPT-6 Sol47.1%measured
49GPT-5.246.9%estimated ± 6.4 pp, low confidence
50Kimi K2.646.7%estimated ± 6.4 pp, low confidence
51Grok 4.346.4%estimated ± 6.4 pp, low confidence
52GPT-5.2-Codex45.0%estimated ± 6.4 pp, low confidence
53MiMo-V2.5-Pro45.0%estimated ± 6.4 pp, low confidence
54Qwen3.7 Plus44.9%estimated ± 6.4 pp, low confidence
55MiMo-V2.6-Flash44.4%estimated ± 6.4 pp, low confidence
56Kimi K2.7 Code44.3%estimated ± 6.4 pp, low confidence
57Mistral Large 444.3%estimated ± 6.4 pp, low confidence
58Apodex 1.143.4%estimated ± 6.4 pp, low confidence
59Apodex 1.1 Mini43.4%estimated ± 6.4 pp, low confidence
60Qwen3.8-27B43.2%estimated ± 6.4 pp, low confidence
61Hy342.8%estimated ± 6.4 pp, low confidence
62Hy3 Preview42.8%estimated ± 6.4 pp, low confidence
63Claude Opus 4.742.6%estimated ± 6.4 pp, low confidence
64Inkling-Small42.6%estimated ± 6.4 pp, low confidence
65Inkling41.2%estimated ± 6.4 pp, low confidence
66Qwen 3.6 Max (preview)40.1%estimated ± 6.4 pp, low confidence
67Kimi K2.539.9%estimated ± 6.4 pp, low confidence
68Kimi K2.5 (Reasoning)39.9%estimated ± 6.4 pp, low confidence
69MiMo-V2-Pro39.6%estimated ± 6.4 pp, low confidence
70Claude Opus 4.5 Thinking39.3%estimated ± 6.4 pp, low confidence
71GLM-5.139.3%estimated ± 6.4 pp, low confidence
72A.X K238.8%estimated ± 6.4 pp, low confidence
73MiniMax M2.738.8%estimated ± 6.4 pp, low confidence
74GLM-538.5%estimated ± 6.4 pp, low confidence
75Solar Pro 438.4%estimated ± 6.4 pp, low confidence
76GPT-5.137.6%estimated ± 6.4 pp, low confidence
77GPT-5 (high)37.6%estimated ± 6.4 pp, low confidence
78Nemotron 3 Ultra37.5%estimated ± 6.4 pp, low confidence
79GPT-5.4 nano37.4%estimated ± 6.4 pp, low confidence
80GPT-5.4 mini37.2%estimated ± 6.4 pp, low confidence
81GLM-5-Turbo36.9%estimated ± 6.4 pp, low confidence
82Qwen3.6 Plus36.9%estimated ± 6.4 pp, low confidence
83GLM-4.736.4%estimated ± 6.4 pp, low confidence
84Grok 435.7%estimated ± 6.4 pp, low confidence
85GPT-5.1-Codex34.6%estimated ± 6.4 pp, low confidence
86GPT-5.1-Codex-Max34.6%estimated ± 6.4 pp, low confidence
87GPT-5 (medium)34.2%estimated ± 6.4 pp, low confidence
88Qwen3.5-122B-A10B34.0%estimated ± 6.4 pp, low confidence
89Qwen3.5-27B32.5%estimated ± 6.4 pp, low confidence
90Ling 3.0 Flash32.3%estimated ± 6.4 pp, low confidence
91Ling 3.0 Flash FP832.3%estimated ± 6.4 pp, low confidence
92Gemma 4 31B32.2%estimated ± 6.4 pp, low confidence
93Qwen3.6-27B31.6%estimated ± 6.4 pp, low confidence
94o3-pro30.9%estimated ± 6.7 pp, low confidence
95Gemini 2.5 Pro30.9%estimated ± 6.4 pp, low confidence
96Qwen3.6-35B-A3B30.6%estimated ± 6.4 pp, low confidence
97MiMo-V2-Omni30.5%estimated ± 6.4 pp, low confidence
98Ling 3.0 Flash VL30.3%estimated ± 6.4 pp, low confidence
99Muse Glimmer 30B30.3%estimated ± 6.4 pp, low confidence
100Step 3.7 Flash29.6%estimated ± 6.4 pp, low confidence
101Qwen3.5-35B-A3B29.2%estimated ± 6.4 pp, low confidence
102Nemotron 3 Super 100B28.9%estimated ± 6.4 pp, low confidence
103o328.1%estimated ± 6.4 pp, low confidence
104Qwen3.5 397B27.7%estimated ± 6.4 pp, low confidence
105Qwen3.5 397B (Reasoning)27.7%estimated ± 6.4 pp, low confidence
106GPT-OSS 120B27.5%estimated ± 6.4 pp, low confidence
107Gemma 4 26B A4B27.1%estimated ± 6.4 pp, low confidence
108Grok 4.1 Fast (Reasoning)27.1%estimated ± 6.4 pp, low confidence
109Claude 4.1 Opus26.9%estimated ± 6.7 pp, low confidence
110Claude Opus 4.626.9%estimated ± 6.4 pp, low confidence
111Grok 4 Fast (Reasoning)26.9%estimated ± 6.4 pp, low confidence
112Gemini 3.5 Flash-Lite26.5%estimated ± 6.4 pp, low confidence
113Quasar 438B26.4%estimated ± 6.4 pp, low confidence
114K-EXAONE 2.026.3%estimated ± 6.4 pp, low confidence
115GLM-5V-Turbo24.4%estimated ± 6.4 pp, low confidence
116DeepSeek-R122.8%estimated ± 6.4 pp, low confidence
117Trinity-Large-Preview22.8%estimated ± 6.4 pp, low confidence
118Trinity-Large-Thinking22.8%estimated ± 6.4 pp, low confidence
119Gemma 4 12B22.6%estimated ± 6.4 pp, low confidence
120Gemini 3 Flash21.7%estimated ± 6.4 pp, low confidence
121DeepSeek V3.1 (Reasoning)20.8%estimated ± 6.4 pp, low confidence
122K-Exaone20.3%estimated ± 6.4 pp, low confidence
123Mistral Medium 3.5 128B20.2%estimated ± 6.4 pp, low confidence
124Claude Sonnet 4.619.5%estimated ± 6.4 pp, low confidence
125Claude Opus 4.519.4%estimated ± 6.4 pp, low confidence
126o1-pro18.8%estimated ± 6.7 pp, low confidence
127Claude 4.1 Opus Thinking18.5%estimated ± 6.4 pp, low confidence
128Command A+17.8%estimated ± 6.4 pp, low confidence
129Qwen3 Max17.6%estimated ± 6.4 pp, low confidence
130Mercury 2.517.5%estimated ± 6.4 pp, low confidence
131o1-preview17.4%estimated ± 6.7 pp, low confidence
132Nemotron 3 Nano 30B17.0%estimated ± 6.4 pp, low confidence
133DeepSeek V3.216.7%estimated ± 6.4 pp, low confidence
134Granite 4.2 30B16.7%estimated ± 6.4 pp, low confidence
135North Mini Code16.6%estimated ± 6.4 pp, low confidence
136GPT-OSS 20B16.4%estimated ± 6.4 pp, low confidence
137Sarvam 105B16.4%estimated ± 6.4 pp, low confidence
138Nemotron 3.5 Lightning 30B A3B NVFP415.9%estimated ± 6.4 pp, low confidence
139Solar Pro 315.5%estimated ± 6.4 pp, low confidence
140Mistral Small 414.9%estimated ± 6.4 pp, low confidence
141Mistral Small 4 (Reasoning)14.9%estimated ± 6.4 pp, low confidence
142Granite 4.2 8B14.6%estimated ± 6.4 pp, low confidence
143Ling 3.0 Tiny14.1%estimated ± 6.4 pp, low confidence
144MiniCPM5-2B13.5%estimated ± 6.4 pp, low confidence
145MiMo-V2-Flash13.1%estimated ± 6.4 pp, low confidence
146Grok Code Fast 112.2%estimated ± 6.4 pp, low confidence
147o3-mini12.1%estimated ± 6.4 pp, low confidence
148Qwen3-Omni-30B-A3B-Thinking11.5%estimated ± 6.4 pp, low confidence
149Sarvam 30B11.5%estimated ± 6.4 pp, low confidence
150Kimi K211.4%estimated ± 6.4 pp, low confidence
151Nemotron Ultra 253B11.4%estimated ± 6.4 pp, low confidence
152GLM-4.5-Air10.8%estimated ± 6.4 pp, low confidence
153o110.8%estimated ± 6.4 pp, low confidence
154LFM2.5-8B-A1B10.6%estimated ± 6.4 pp, low confidence
155Celeris-110.5%estimated ± 6.4 pp, low confidence
156DeepSeek V3.110.3%estimated ± 6.4 pp, low confidence
157Granite 4.2 3B10.2%estimated ± 6.4 pp, low confidence
158Granite-4.0-H-350M9.9%estimated ± 6.4 pp, low confidence
159Ling 2.6 Flash9.7%estimated ± 6.4 pp, low confidence
160LFM2.5-2.6B9.6%estimated ± 6.4 pp, low confidence
161Exaone 4.0 1.2B8.9%estimated ± 6.4 pp, low confidence
162GLM-4.68.6%estimated ± 6.4 pp, low confidence
163Granite-4.0-350M8.6%estimated ± 6.4 pp, low confidence
164Grok 4.1 Fast8.0%estimated ± 6.4 pp, low confidence
165LFM2.5-VL-1.6B-Extract8.0%estimated ± 6.4 pp, low confidence
166Exaone 4.0 32B7.8%estimated ± 6.4 pp, low confidence
167GPT-4.1 mini7.8%estimated ± 6.4 pp, low confidence
168Granite-4.0-H-1B7.8%estimated ± 6.4 pp, low confidence
169Phi-4 Multimodal Instruct7.8%estimated ± 6.4 pp, low confidence
170Llama 4 Maverick7.7%estimated ± 6.4 pp, low confidence
171Gemma 4 E2B7.5%estimated ± 6.4 pp, low confidence
172Nemotron 3 Nano Omni 30B A3B7.5%estimated ± 6.4 pp, low confidence
173DeepSeek V3 03247.4%estimated ± 6.4 pp, low confidence
174Gemini 2.5 Flash7.4%estimated ± 6.4 pp, low confidence
175DeepSeek R1 Distill Qwen 32B7.2%estimated ± 6.4 pp, low confidence
176Gemini 1.5 Pro7.2%estimated ± 6.4 pp, low confidence
177Qwen3-Omni-30B-A3B-Instruct7.2%estimated ± 6.4 pp, low confidence
178Gemma 3 27B6.9%estimated ± 6.4 pp, low confidence
179Claude 4 Sonnet6.8%estimated ± 6.4 pp, low confidence
180Gemini 1.0 Pro6.6%estimated ± 6.4 pp, low confidence
181GPT-4.16.6%estimated ± 6.4 pp, low confidence
182GPT-4o mini6.6%estimated ± 6.4 pp, low confidence
183Mistral Large 36.6%estimated ± 6.4 pp, low confidence
184Claude 3 Haiku6.5%estimated ± 6.4 pp, low confidence
185Mistral Medium 36.5%estimated ± 6.4 pp, low confidence
186Llama 3.1 405B6.3%estimated ± 6.4 pp, low confidence
187Gemma 4 E4B6.0%estimated ± 6.4 pp, low confidence
188GPT-4.1 nano6.0%estimated ± 6.4 pp, low confidence
189Llama 4 Scout6.0%estimated ± 6.4 pp, low confidence
190Phi-46.0%estimated ± 6.4 pp, low confidence
191Solar Pro 25.8%estimated ± 6.4 pp, low confidence
192Ultravox v0.6 Llama 3.3 70B5.7%estimated ± 6.4 pp, low confidence
193Qwen2.5 Coder 32B Instruct5.5%estimated ± 6.4 pp, low confidence
194Mistral Large 25.2%estimated ± 6.4 pp, low confidence
195Nova Pro5.1%estimated ± 6.4 pp, low confidence
196GPT-4 Turbo4.9%estimated ± 6.4 pp, low confidence
197DeepSeek V34.6%estimated ± 6.4 pp, low confidence
198Claude 3 Opus4.5%estimated ± 6.4 pp, low confidence
199GPT-4o3.8%estimated ± 6.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General