benchgap
Calibration

AA-HLE → HealthBench Professional (raw)

HealthBench Professional (raw) is estimated from AA-HLE with a Michaelis–Menten + offset curve fitted on 7 models measured on both: y = 0.1510 + 2.0000·x / (1.40432 + x), R² = 0.66, cross-validated error 5.0 pp. It is used for 188 estimates.

Estimated modelAA-HLEHealthBench Professional (raw)Source
A.X K229.6%49.9%estimated ± 5.0 pp, low confidence
Apodex 1.134.1%54.2%estimated ± 5.0 pp, low confidence
Apodex 1.1 Mini34.1%54.2%estimated ± 5.0 pp, low confidence
Celeris-16.8%24.3%estimated ± 5.0 pp, low confidence
Claude 3 Haiku4.1%20.8%estimated ± 5.0 pp, low confidence
Claude 3 Opus2.8%19.0%estimated ± 5.0 pp, low confidence
Claude 4.1 Opus Thinking12.5%31.4%estimated ± 5.0 pp, low confidence
Claude 4 Sonnet4.3%21.0%estimated ± 5.0 pp, low confidence
Claude Fable 555.5%71.7%estimated ± 5.0 pp, medium confidence
Claude Fable 5.159.1%74.3%estimated ± 5.0 pp, medium confidence
Claude Haiku 5.544.4%63.1%estimated ± 5.0 pp, medium confidence
Claude Opus 4.513.2%32.3%estimated ± 5.0 pp, low confidence
Claude Opus 4.5 Thinking30.1%50.4%estimated ± 5.0 pp, low confidence
Claude Opus 4.619.1%39.0%estimated ± 5.0 pp, low confidence
Claude Opus 4.6 (Adaptive)39.9%59.3%estimated ± 5.0 pp, medium confidence
Claude Opus 4.733.3%53.4%estimated ± 5.0 pp, low confidence
Claude Opus 4.7 (Adaptive)42.3%61.4%estimated ± 5.0 pp, medium confidence
Claude Opus 4.848.7%66.6%estimated ± 5.0 pp, medium confidence
Claude Sonnet 4.613.3%32.4%estimated ± 5.0 pp, low confidence
Claude Sonnet 541.3%60.5%estimated ± 5.0 pp, medium confidence
Command A+12.0%30.8%estimated ± 5.0 pp, low confidence
DeepSeek-R115.8%35.3%estimated ± 5.0 pp, low confidence
DeepSeek R1 Distill Qwen 32B4.6%21.4%estimated ± 5.0 pp, low confidence
DeepSeek V32.9%19.1%estimated ± 5.0 pp, low confidence
DeepSeek V3 03244.7%21.6%estimated ± 5.0 pp, low confidence
DeepSeek V3.16.7%24.2%estimated ± 5.0 pp, low confidence
DeepSeek V3.1 (Reasoning)14.3%33.6%estimated ± 5.0 pp, low confidence
DeepSeek V3.211.2%29.9%estimated ± 5.0 pp, low confidence
DeepSeek V4.1 Flash39.2%58.7%estimated ± 5.0 pp, medium confidence
DeepSeek V4 Flash 073138.6%58.2%estimated ± 5.0 pp, medium confidence
DeepSeek V4 Pro 081341.0%60.3%estimated ± 5.0 pp, medium confidence
Exaone 4.0 1.2B5.7%22.9%estimated ± 5.0 pp, low confidence
Exaone 4.0 32B5.0%22.0%estimated ± 5.0 pp, low confidence
Gemini 1.0 Pro4.2%20.9%estimated ± 5.0 pp, low confidence
Gemini 1.5 Pro4.6%21.4%estimated ± 5.0 pp, low confidence
Gemini 2.5 Flash4.7%21.6%estimated ± 5.0 pp, low confidence
Gemini 2.5 Pro22.5%42.7%estimated ± 5.0 pp, low confidence
Gemini 3.1 Pro47.0%65.2%estimated ± 5.0 pp, medium confidence
Gemini 3.5 Flash42.7%61.7%estimated ± 5.0 pp, medium confidence
Gemini 3.5 Flash-Lite18.8%38.7%estimated ± 5.0 pp, low confidence
Gemini 3.6 Flash40.8%60.1%estimated ± 5.0 pp, medium confidence
Gemini 3.7 Flash47.9%66.0%estimated ± 5.0 pp, medium confidence
Gemini 3.8 Flash47.8%65.9%estimated ± 5.0 pp, medium confidence
Gemini 3 Flash15.0%34.4%estimated ± 5.0 pp, low confidence
Gemini 3 Pro39.7%59.2%estimated ± 5.0 pp, medium confidence
Gemini 4 Argon57.1%72.9%estimated ± 5.0 pp, medium confidence
Gemma 3 27B4.4%21.2%estimated ± 5.0 pp, low confidence
Gemma 4 12B15.7%35.2%estimated ± 5.0 pp, low confidence
Gemma 4 26B A4B19.3%39.3%estimated ± 5.0 pp, low confidence
Gemma 4 31B23.6%43.9%estimated ± 5.0 pp, low confidence
Gemma 4 E2B4.8%21.7%estimated ± 5.0 pp, low confidence
Gemma 4 E4B3.8%20.4%estimated ± 5.0 pp, low confidence
GLM-4.5-Air7.0%24.6%estimated ± 5.0 pp, low confidence
GLM-4.65.5%22.6%estimated ± 5.0 pp, low confidence
GLM-4.727.4%47.7%estimated ± 5.0 pp, low confidence
GLM-529.3%49.6%estimated ± 5.0 pp, low confidence
GLM-5.130.1%50.4%estimated ± 5.0 pp, low confidence
GLM-5.241.1%60.4%estimated ± 5.0 pp, medium confidence
GLM-5.342.3%61.4%estimated ± 5.0 pp, medium confidence
GLM-5.3-Flash39.9%59.3%estimated ± 5.0 pp, medium confidence
GLM-5-Turbo27.8%48.1%estimated ± 5.0 pp, low confidence
GLM-5V-Turbo17.1%36.8%estimated ± 5.0 pp, low confidence
GPT-4.14.2%20.9%estimated ± 5.0 pp, low confidence
GPT-4.1 mini5.0%22.0%estimated ± 5.0 pp, low confidence
GPT-4.1 nano3.8%20.4%estimated ± 5.0 pp, low confidence
GPT-4 Turbo3.1%19.4%estimated ± 5.0 pp, low confidence
GPT-4o2.4%18.5%estimated ± 5.0 pp, low confidence
GPT-4o mini4.2%20.9%estimated ± 5.0 pp, low confidence
GPT-5.128.5%48.8%estimated ± 5.0 pp, low confidence
GPT-5.1-Codex25.7%46.0%estimated ± 5.0 pp, low confidence
GPT-5.1-Codex-Max25.7%46.0%estimated ± 5.0 pp, low confidence
GPT-5.237.7%57.4%estimated ± 5.0 pp, low confidence
GPT-5.2-Codex35.7%55.6%estimated ± 5.0 pp, low confidence
GPT-5.3 Codex42.5%61.6%estimated ± 5.0 pp, medium confidence
GPT-5.443.7%62.6%estimated ± 5.0 pp, medium confidence
GPT-5.4 mini28.1%48.4%estimated ± 5.0 pp, low confidence
GPT-5.4 nano28.3%48.6%estimated ± 5.0 pp, low confidence
GPT-5.545.8%64.3%estimated ± 5.0 pp, medium confidence
GPT-5.6 Luna39.5%59.0%estimated ± 5.0 pp, medium confidence
GPT-5.6 Sol49.5%67.2%estimated ± 5.0 pp, medium confidence
GPT-5.6 Terra42.9%61.9%estimated ± 5.0 pp, medium confidence
GPT-5 (high)28.5%48.8%estimated ± 5.0 pp, low confidence
GPT-5 (medium)25.4%45.7%estimated ± 5.0 pp, low confidence
GPT-OSS 120B19.6%39.6%estimated ± 5.0 pp, low confidence
GPT-OSS 20B11.0%29.6%estimated ± 5.0 pp, low confidence
Granite-4.0-350M5.5%22.6%estimated ± 5.0 pp, low confidence
Granite-4.0-H-1B5.0%22.0%estimated ± 5.0 pp, low confidence
Granite-4.0-H-350M6.4%23.8%estimated ± 5.0 pp, low confidence
Granite 4.2 30B11.2%29.9%estimated ± 5.0 pp, low confidence
Granite 4.2 3B6.6%24.1%estimated ± 5.0 pp, low confidence
Granite 4.2 8B9.7%28.0%estimated ± 5.0 pp, low confidence
Grok 426.7%47.0%estimated ± 5.0 pp, low confidence
Grok 4.1 Fast5.1%22.1%estimated ± 5.0 pp, low confidence
Grok 4.1 Fast (Reasoning)19.3%39.3%estimated ± 5.0 pp, low confidence
Grok 4.337.2%57.0%estimated ± 5.0 pp, low confidence
Grok 4.542.7%61.7%estimated ± 5.0 pp, medium confidence
Grok 4.642.9%61.9%estimated ± 5.0 pp, medium confidence
Grok 4.743.1%62.1%estimated ± 5.0 pp, medium confidence
Grok 4 Fast (Reasoning)19.1%39.0%estimated ± 5.0 pp, low confidence
Grok Code Fast 18.0%25.9%estimated ± 5.0 pp, low confidence
Hy333.5%53.6%estimated ± 5.0 pp, low confidence
Hy3 Preview33.5%53.6%estimated ± 5.0 pp, low confidence
Inkling31.9%52.1%estimated ± 5.0 pp, low confidence
Inkling-Small33.3%53.4%estimated ± 5.0 pp, low confidence
K-Exaone13.9%33.1%estimated ± 5.0 pp, low confidence
K-EXAONE 2.018.6%38.5%estimated ± 5.0 pp, low confidence
Kimi K2.637.5%57.2%estimated ± 5.0 pp, low confidence
Kimi K27.4%25.1%estimated ± 5.0 pp, low confidence
Kimi K2.530.7%51.0%estimated ± 5.0 pp, low confidence
Kimi K2.5 (Reasoning)30.7%51.0%estimated ± 5.0 pp, low confidence
Kimi K2.7 Code35.0%55.0%estimated ± 5.0 pp, low confidence
Kimi K346.9%65.2%estimated ± 5.0 pp, medium confidence
LFM2.5-2.6B6.2%23.6%estimated ± 5.0 pp, low confidence
LFM2.5-8B-A1B6.9%24.5%estimated ± 5.0 pp, low confidence
LFM2.5-VL-1.6B-Extract5.1%22.1%estimated ± 5.0 pp, low confidence
Ling 2.6 Flash6.3%23.7%estimated ± 5.0 pp, low confidence
Ling 3.0 Flash23.7%44.0%estimated ± 5.0 pp, low confidence
Ling 3.0 Flash FP823.7%44.0%estimated ± 5.0 pp, low confidence
Ling 3.0 Flash VL22.0%42.2%estimated ± 5.0 pp, low confidence
Ling 3.0 Tiny9.3%27.5%estimated ± 5.0 pp, low confidence
Ling 3.1 Flash39.4%58.9%estimated ± 5.0 pp, medium confidence
Llama 3.1 405B4.0%20.6%estimated ± 5.0 pp, low confidence
Llama 4 Maverick4.9%21.8%estimated ± 5.0 pp, low confidence
Llama 4 Scout3.8%20.4%estimated ± 5.0 pp, low confidence
Mercury 2.511.8%30.6%estimated ± 5.0 pp, low confidence
MiMo-V2.5-Pro35.7%55.6%estimated ± 5.0 pp, low confidence
MiMo-V2.6-Flash35.1%55.1%estimated ± 5.0 pp, low confidence
MiMo-V2.6-Pro49.4%67.1%estimated ± 5.0 pp, medium confidence
MiMo-V2-Flash8.6%26.6%estimated ± 5.0 pp, low confidence
MiMo-V2-Omni22.1%42.3%estimated ± 5.0 pp, low confidence
MiMo-V2-Pro30.4%50.7%estimated ± 5.0 pp, low confidence
MiniCPM5-2B8.9%27.0%estimated ± 5.0 pp, low confidence
MiniMax M2.729.6%49.9%estimated ± 5.0 pp, low confidence
MiniMax M339.0%58.6%estimated ± 5.0 pp, medium confidence
Mistral Large 23.3%19.7%estimated ± 5.0 pp, low confidence
Mistral Large 34.2%20.9%estimated ± 5.0 pp, low confidence
Mistral Large 435.0%55.0%estimated ± 5.0 pp, low confidence
Mistral Medium 34.1%20.8%estimated ± 5.0 pp, low confidence
Mistral Medium 3.5 128B13.8%33.0%estimated ± 5.0 pp, low confidence
Mistral Small 49.9%28.3%estimated ± 5.0 pp, low confidence
Mistral Small 4 (Reasoning)9.9%28.3%estimated ± 5.0 pp, low confidence
Muse Glimmer 30B22.0%42.2%estimated ± 5.0 pp, low confidence
Muse Spark40.7%60.0%estimated ± 5.0 pp, medium confidence
Muse Spark 1.146.2%64.6%estimated ± 5.0 pp, medium confidence
Muse Spark 1.245.5%64.0%estimated ± 5.0 pp, medium confidence
Muse Spark 1.348.7%66.6%estimated ± 5.0 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP410.6%29.1%estimated ± 5.0 pp, low confidence
Nemotron 3 Nano 30B11.4%30.1%estimated ± 5.0 pp, low confidence
Nemotron 3 Nano Omni 30B A3B4.8%21.7%estimated ± 5.0 pp, low confidence
Nemotron 3 Super 100B20.8%40.9%estimated ± 5.0 pp, low confidence
Nemotron 3 Ultra28.4%48.7%estimated ± 5.0 pp, low confidence
Nemotron Ultra 253B7.4%25.1%estimated ± 5.0 pp, low confidence
North Mini Code11.1%29.7%estimated ± 5.0 pp, low confidence
Nova Pro3.2%19.6%estimated ± 5.0 pp, low confidence
o17.0%24.6%estimated ± 5.0 pp, low confidence
o320.1%40.1%estimated ± 5.0 pp, low confidence
o3-mini7.9%25.7%estimated ± 5.0 pp, low confidence
Phi-43.8%20.4%estimated ± 5.0 pp, low confidence
Phi-4 Multimodal Instruct5.0%22.0%estimated ± 5.0 pp, low confidence
Quasar 438B18.7%38.6%estimated ± 5.0 pp, low confidence
Qwen2.5 Coder 32B Instruct3.5%20.0%estimated ± 5.0 pp, low confidence
Qwen3.5-122B-A10B25.2%45.5%estimated ± 5.0 pp, low confidence
Qwen3.5-27B23.9%44.2%estimated ± 5.0 pp, low confidence
Qwen3.5-35B-A3B21.0%41.1%estimated ± 5.0 pp, low confidence
Qwen3.5 397B19.8%39.8%estimated ± 5.0 pp, low confidence
Qwen3.5 397B (Reasoning)19.8%39.8%estimated ± 5.0 pp, low confidence
Qwen3.6-27B23.1%43.3%estimated ± 5.0 pp, low confidence
Qwen3.6-35B-A3B22.2%42.4%estimated ± 5.0 pp, low confidence
Qwen 3.6 Max (preview)30.8%51.1%estimated ± 5.0 pp, low confidence
Qwen3.6 Plus27.8%48.1%estimated ± 5.0 pp, low confidence
Qwen3.7 Max40.5%59.9%estimated ± 5.0 pp, medium confidence
Qwen3.7 Plus35.6%55.5%estimated ± 5.0 pp, low confidence
Qwen3.8-27B33.9%54.0%estimated ± 5.0 pp, low confidence
Qwen3.8-Flash-Next38.0%57.7%estimated ± 5.0 pp, low confidence
Qwen3.8 Max Preview43.1%62.1%estimated ± 5.0 pp, medium confidence
Qwen3 Max11.9%30.7%estimated ± 5.0 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct4.6%21.4%estimated ± 5.0 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking7.5%25.2%estimated ± 5.0 pp, low confidence
Sarvam 105B11.0%29.6%estimated ± 5.0 pp, low confidence
Sarvam 30B7.5%25.2%estimated ± 5.0 pp, low confidence
Solar Pro 23.7%20.2%estimated ± 5.0 pp, low confidence
Solar Pro 310.3%28.8%estimated ± 5.0 pp, low confidence
Solar Pro 429.2%49.5%estimated ± 5.0 pp, low confidence
Step 3.7 Flash21.4%41.5%estimated ± 5.0 pp, low confidence
Step 5 Preview46.5%64.8%estimated ± 5.0 pp, medium confidence
Trinity-Large-Preview15.8%35.3%estimated ± 5.0 pp, low confidence
Trinity-Large-Thinking15.8%35.3%estimated ± 5.0 pp, low confidence
Ultravox v0.6 Llama 3.3 70B3.6%20.1%estimated ± 5.0 pp, low confidence