benchgap
Knowledge & reasoning

HealthBench (length-adjusted) leaderboard

As of 2026-10-07, the highest measured score on HealthBench (length-adjusted) is 65.4% by Claude Sonnet 5.5. 196 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.565.4%measured
2Claude Opus 5.560.6%measured
3Ling 3.1 Flash60.6%estimated ± 2.2 pp, medium confidence
4Gemini 4 Argon59.5%estimated ± 3.7 pp, medium confidence
5GPT-6.1 Sol58.5%measured
6Claude Fable 558.3%estimated ± 2.5 pp, low confidence
7GPT-6 Astra58.3%measured
8Claude Fable 5.158.0%estimated ± 2.5 pp, low confidence
9Claude Opus 557.8%measured
10Gemini 3.7 Flash57.2%estimated ± 2.5 pp, low confidence
11GPT-5.5 Pro57.0%estimated ± 2.5 pp, low confidence
12GPT-5.4 Pro56.8%estimated ± 2.5 pp, low confidence
13Kimi K356.8%estimated ± 2.5 pp, low confidence
14Claude Opus 4.6 (Adaptive)56.3%estimated ± 2.5 pp, low confidence
15Gemini 3.5 Flash56.1%estimated ± 2.5 pp, low confidence
16Gemini 3.6 Flash55.6%estimated ± 2.5 pp, low confidence
17DeepSeek V4 Pro 081355.1%estimated ± 2.5 pp, low confidence
18DeepSeek V4 Flash 073154.7%estimated ± 2.5 pp, low confidence
19GPT-6 Luna54.5%measured
20Muse Spark 1.354.4%estimated ± 3.7 pp, medium confidence
21MiMo-V2.6-Pro54.1%estimated ± 3.7 pp, medium confidence
22Qwen3.8 Max Preview54.0%estimated ± 3.7 pp, medium confidence
23GLM-5.354.0%estimated ± 3.7 pp, medium confidence
24Step 5 Preview53.9%estimated ± 3.7 pp, medium confidence
25Claude Haiku 5.553.9%estimated ± 3.7 pp, medium confidence
26GLM-5.3-Flash53.9%estimated ± 3.7 pp, medium confidence
27Hy3 Preview53.9%estimated ± 3.7 pp, medium confidence
28Qwen3.8-Flash-Next53.9%estimated ± 3.7 pp, medium confidence
29Muse Spark 1.253.9%estimated ± 3.7 pp, medium confidence
30DeepSeek V4.1 Flash53.9%estimated ± 3.7 pp, medium confidence
31Mistral Large 453.9%estimated ± 3.7 pp, medium confidence
32Claude Sonnet 553.9%estimated ± 3.7 pp, medium confidence
33Grok 4.353.9%estimated ± 3.7 pp, low confidence
34MiMo-V2.6-Flash53.9%estimated ± 3.7 pp, low confidence
35GLM-5.253.9%estimated ± 3.7 pp, low confidence
36GPT-5.3 Codex53.9%estimated ± 3.7 pp, low confidence
37Qwen3.8-27B53.9%estimated ± 3.7 pp, low confidence
38A.X K253.9%estimated ± 3.7 pp, low confidence
39Apodex 1.153.9%estimated ± 3.7 pp, low confidence
40Apodex 1.1 Mini53.9%estimated ± 3.7 pp, low confidence
41Celeris-153.9%estimated ± 3.7 pp, low confidence
42Claude 3 Haiku53.9%estimated ± 3.7 pp, low confidence
43Claude 3 Opus53.9%estimated ± 3.7 pp, low confidence
44Claude 4.1 Opus53.9%estimated ± 3.7 pp, low confidence
45Claude 4.1 Opus Thinking53.9%estimated ± 3.7 pp, low confidence
46Claude 4 Sonnet53.9%estimated ± 3.7 pp, low confidence
47Claude Opus 4.553.9%estimated ± 3.7 pp, low confidence
48Claude Opus 4.653.9%estimated ± 3.7 pp, low confidence
49Claude Opus 4.753.9%estimated ± 3.7 pp, low confidence
50Command A+53.9%estimated ± 3.7 pp, low confidence
51DeepSeek-R153.9%estimated ± 3.7 pp, low confidence
52DeepSeek R1 Distill Qwen 32B53.9%estimated ± 3.7 pp, low confidence
53DeepSeek V353.9%estimated ± 3.7 pp, low confidence
54DeepSeek V3 032453.9%estimated ± 3.7 pp, low confidence
55DeepSeek V3.153.9%estimated ± 3.7 pp, low confidence
56DeepSeek V3.1 (Reasoning)53.9%estimated ± 3.7 pp, low confidence
57DeepSeek V3.253.9%estimated ± 3.7 pp, low confidence
58Exaone 4.0 1.2B53.9%estimated ± 3.7 pp, low confidence
59Exaone 4.0 32B53.9%estimated ± 3.7 pp, low confidence
60Gemini 1.0 Pro53.9%estimated ± 3.7 pp, low confidence
61Gemini 1.5 Pro53.9%estimated ± 3.7 pp, low confidence
62Gemini 2.5 Flash53.9%estimated ± 3.7 pp, low confidence
63Gemini 2.5 Pro53.9%estimated ± 3.7 pp, low confidence
64Gemini 3.5 Flash-Lite53.9%estimated ± 3.7 pp, low confidence
65Gemini 3 Flash53.9%estimated ± 3.7 pp, low confidence
66Gemma 3 27B53.9%estimated ± 3.7 pp, low confidence
67Gemma 4 12B53.9%estimated ± 3.7 pp, low confidence
68Gemma 4 26B A4B53.9%estimated ± 3.7 pp, low confidence
69Gemma 4 31B53.9%estimated ± 3.7 pp, low confidence
70Gemma 4 E2B53.9%estimated ± 3.7 pp, low confidence
71Gemma 4 E4B53.9%estimated ± 3.7 pp, low confidence
72GLM-4.5-Air53.9%estimated ± 3.7 pp, low confidence
73GLM-4.653.9%estimated ± 3.7 pp, low confidence
74GLM-4.753.9%estimated ± 3.7 pp, low confidence
75GLM-553.9%estimated ± 3.7 pp, low confidence
76GLM-5.153.9%estimated ± 3.7 pp, low confidence
77GLM-5-Turbo53.9%estimated ± 3.7 pp, low confidence
78GLM-5V-Turbo53.9%estimated ± 3.7 pp, low confidence
79GPT-4.153.9%estimated ± 3.7 pp, low confidence
80GPT-4.1 mini53.9%estimated ± 3.7 pp, low confidence
81GPT-4.1 nano53.9%estimated ± 3.7 pp, low confidence
82GPT-4 Turbo53.9%estimated ± 3.7 pp, low confidence
83GPT-4o53.9%estimated ± 3.7 pp, low confidence
84GPT-4o mini53.9%estimated ± 3.7 pp, low confidence
85GPT-5.1-Codex53.9%estimated ± 3.7 pp, low confidence
86GPT-5.1-Codex-Max53.9%estimated ± 3.7 pp, low confidence
87GPT-5.2-Codex53.9%estimated ± 3.7 pp, low confidence
88GPT-5 (high)53.9%estimated ± 3.7 pp, low confidence
89GPT-5 (medium)53.9%estimated ± 3.7 pp, low confidence
90GPT-OSS 120B53.9%estimated ± 3.7 pp, low confidence
91GPT-OSS 20B53.9%estimated ± 3.7 pp, low confidence
92Granite-4.0-350M53.9%estimated ± 3.7 pp, low confidence
93Granite-4.0-H-1B53.9%estimated ± 3.7 pp, low confidence
94Granite-4.0-H-350M53.9%estimated ± 3.7 pp, low confidence
95Granite 4.2 30B53.9%estimated ± 3.7 pp, low confidence
96Granite 4.2 3B53.9%estimated ± 3.7 pp, low confidence
97Granite 4.2 8B53.9%estimated ± 3.7 pp, low confidence
98Grok 453.9%estimated ± 3.7 pp, low confidence
99Grok 4.1 Fast53.9%estimated ± 3.7 pp, low confidence
100Grok 4.1 Fast (Reasoning)53.9%estimated ± 3.7 pp, low confidence
101Grok 4 Fast (Reasoning)53.9%estimated ± 3.7 pp, low confidence
102Grok Code Fast 153.9%estimated ± 3.7 pp, low confidence
103Hy353.9%estimated ± 3.7 pp, low confidence
104Inkling53.9%estimated ± 3.7 pp, low confidence
105K-Exaone53.9%estimated ± 3.7 pp, low confidence
106K-EXAONE 2.053.9%estimated ± 3.7 pp, low confidence
107Kimi K2.653.9%estimated ± 3.7 pp, low confidence
108Kimi K253.9%estimated ± 3.7 pp, low confidence
109Kimi K2.553.9%estimated ± 3.7 pp, low confidence
110Kimi K2.5 (Reasoning)53.9%estimated ± 3.7 pp, low confidence
111Kimi K2.7 Code53.9%estimated ± 3.7 pp, low confidence
112LFM2.5-2.6B53.9%estimated ± 3.7 pp, low confidence
113LFM2.5-8B-A1B53.9%estimated ± 3.7 pp, low confidence
114LFM2.5-VL-1.6B-Extract53.9%estimated ± 3.7 pp, low confidence
115Ling 2.6 Flash53.9%estimated ± 3.7 pp, low confidence
116Ling 3.0 Flash53.9%estimated ± 3.7 pp, low confidence
117Ling 3.0 Flash FP853.9%estimated ± 3.7 pp, low confidence
118Ling 3.0 Flash VL53.9%estimated ± 3.7 pp, low confidence
119Ling 3.0 Tiny53.9%estimated ± 3.7 pp, low confidence
120Llama 3.1 405B53.9%estimated ± 3.7 pp, low confidence
121Llama 4 Maverick53.9%estimated ± 3.7 pp, low confidence
122Llama 4 Scout53.9%estimated ± 3.7 pp, low confidence
123Mercury 2.553.9%estimated ± 3.7 pp, low confidence
124MiMo-V2.5-Pro53.9%estimated ± 3.7 pp, low confidence
125MiMo-V2-Flash53.9%estimated ± 3.7 pp, low confidence
126MiMo-V2-Omni53.9%estimated ± 3.7 pp, low confidence
127MiMo-V2-Pro53.9%estimated ± 3.7 pp, low confidence
128MiniCPM5-2B53.9%estimated ± 3.7 pp, low confidence
129MiniMax M2.753.9%estimated ± 3.7 pp, low confidence
130MiniMax M353.9%estimated ± 3.7 pp, low confidence
131Mistral Large 253.9%estimated ± 3.7 pp, low confidence
132Mistral Large 353.9%estimated ± 3.7 pp, low confidence
133Mistral Medium 353.9%estimated ± 3.7 pp, low confidence
134Mistral Medium 3.5 128B53.9%estimated ± 3.7 pp, low confidence
135Mistral Small 453.9%estimated ± 3.7 pp, low confidence
136Mistral Small 4 (Reasoning)53.9%estimated ± 3.7 pp, low confidence
137Muse Glimmer 30B53.9%estimated ± 3.7 pp, low confidence
138Muse Spark53.9%estimated ± 3.7 pp, low confidence
139Nemotron 3.5 Lightning 30B A3B NVFP453.9%estimated ± 3.7 pp, low confidence
140Nemotron 3 Nano 30B53.9%estimated ± 3.7 pp, low confidence
141Nemotron 3 Nano Omni 30B A3B53.9%estimated ± 3.7 pp, low confidence
142Nemotron 3 Super 100B53.9%estimated ± 3.7 pp, low confidence
143Nemotron 3 Ultra53.9%estimated ± 3.7 pp, low confidence
144Nemotron Ultra 253B53.9%estimated ± 3.7 pp, low confidence
145North Mini Code53.9%estimated ± 3.7 pp, low confidence
146Nova Pro53.9%estimated ± 3.7 pp, low confidence
147o153.9%estimated ± 3.7 pp, low confidence
148o1-preview53.9%estimated ± 3.7 pp, low confidence
149o1-pro53.9%estimated ± 3.7 pp, low confidence
150o353.9%estimated ± 3.7 pp, low confidence
151o3-mini53.9%estimated ± 3.7 pp, low confidence
152o3-pro53.9%estimated ± 3.7 pp, low confidence
153Phi-453.9%estimated ± 3.7 pp, low confidence
154Phi-4 Multimodal Instruct53.9%estimated ± 3.7 pp, low confidence
155Quasar 438B53.9%estimated ± 3.7 pp, low confidence
156Qwen2.5 Coder 32B Instruct53.9%estimated ± 3.7 pp, low confidence
157Qwen3.5-122B-A10B53.9%estimated ± 3.7 pp, low confidence
158Qwen3.5-27B53.9%estimated ± 3.7 pp, low confidence
159Qwen3.5-35B-A3B53.9%estimated ± 3.7 pp, low confidence
160Qwen3.5 397B53.9%estimated ± 3.7 pp, low confidence
161Qwen3.5 397B (Reasoning)53.9%estimated ± 3.7 pp, low confidence
162Qwen3.6-27B53.9%estimated ± 3.7 pp, low confidence
163Qwen3.6-35B-A3B53.9%estimated ± 3.7 pp, low confidence
164Qwen 3.6 Max (preview)53.9%estimated ± 3.7 pp, low confidence
165Qwen3.6 Plus53.9%estimated ± 3.7 pp, low confidence
166Qwen3.7 Max53.9%estimated ± 3.7 pp, low confidence
167Qwen3.7 Plus53.9%estimated ± 3.7 pp, low confidence
168Qwen3 Max53.9%estimated ± 3.7 pp, low confidence
169Qwen3-Omni-30B-A3B-Instruct53.9%estimated ± 3.7 pp, low confidence
170Qwen3-Omni-30B-A3B-Thinking53.9%estimated ± 3.7 pp, low confidence
171Sarvam 105B53.9%estimated ± 3.7 pp, low confidence
172Sarvam 30B53.9%estimated ± 3.7 pp, low confidence
173Solar Pro 253.9%estimated ± 3.7 pp, low confidence
174Solar Pro 353.9%estimated ± 3.7 pp, low confidence
175Solar Pro 453.9%estimated ± 3.7 pp, low confidence
176Step 3.7 Flash53.9%estimated ± 3.7 pp, low confidence
177Trinity-Large-Preview53.9%estimated ± 3.7 pp, low confidence
178Trinity-Large-Thinking53.9%estimated ± 3.7 pp, low confidence
179Ultravox v0.6 Llama 3.3 70B53.9%estimated ± 3.7 pp, low confidence
180Gemini 3.8 Flash53.9%estimated ± 0.9 pp, medium confidence
181GPT-5.6 Sol53.9%estimated ± 0.9 pp, medium confidence
182Claude Opus 4.7 (Adaptive)53.8%estimated ± 0.9 pp, medium confidence
183Claude Opus 4.853.8%estimated ± 0.9 pp, medium confidence
184Gemini 3.1 Pro53.8%estimated ± 0.9 pp, medium confidence
185GPT-5.453.8%estimated ± 0.9 pp, medium confidence
186GPT-5.553.8%estimated ± 0.9 pp, medium confidence
187GPT-5.6 Luna53.8%estimated ± 0.9 pp, medium confidence
188GPT-5.6 Terra53.8%estimated ± 0.9 pp, medium confidence
189Grok 4.2053.8%estimated ± 0.9 pp, low confidence
190Grok 4.553.8%estimated ± 0.9 pp, medium confidence
191Grok 4.653.8%estimated ± 0.9 pp, medium confidence
192GPT-5.253.6%estimated ± 2.5 pp, low confidence
193Claude Sonnet 4.653.5%estimated ± 2.5 pp, low confidence
194GPT-6 Sol53.2%measured
195Inkling-Small52.7%estimated ± 2.5 pp, low confidence
196Muse Spark 1.152.6%estimated ± 2.2 pp, low confidence
197Claude Opus 4.5 Thinking51.1%estimated ± 2.5 pp, low confidence
198Grok 4.749.4%estimated ± 2.2 pp, low confidence
199Gemini 3 Pro48.9%estimated ± 2.5 pp, low confidence
200GPT-5.147.9%estimated ± 2.5 pp, low confidence
201GPT-5.4 mini43.7%estimated ± 2.5 pp, low confidence
202Claude Sonnet 4.5 Thinking43.6%estimated ± 2.5 pp, low confidence
203GPT-5.4 nano37.4%estimated ± 2.5 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General