benchgap
Calibration

Artificial Analysis Intelligence Index → HealthBench (length-adjusted)

HealthBench (length-adjusted) is estimated from Artificial Analysis Intelligence Index with a offset logistic curve fitted on 7 models measured on both: y = 0.5390 + (0.6296 − 0.5390) / (1 + exp(−73.67·(x − 0.5188))), R² = 0.81, cross-validated error 3.7 pp. It is used for 161 estimates.

Estimated modelArtificial Analysis Intelligence IndexHealthBench (length-adjusted)Source
A.X K221.5%53.9%estimated ± 3.7 pp, low confidence
Apodex 1.126.4%53.9%estimated ± 3.7 pp, low confidence
Apodex 1.1 Mini26.4%53.9%estimated ± 3.7 pp, low confidence
Celeris-16.4%53.9%estimated ± 3.7 pp, low confidence
Claude 3 Haiku5.6%53.9%estimated ± 3.7 pp, low confidence
Claude 3 Opus8.7%53.9%estimated ± 3.7 pp, low confidence
Claude 4.1 Opus18.6%53.9%estimated ± 3.7 pp, low confidence
Claude 4.1 Opus Thinking22.9%53.9%estimated ± 3.7 pp, low confidence
Claude 4 Sonnet16.6%53.9%estimated ± 3.7 pp, low confidence
Claude Haiku 5.543.4%53.9%estimated ± 3.7 pp, medium confidence
Claude Opus 4.523.7%53.9%estimated ± 3.7 pp, low confidence
Claude Opus 4.626.4%53.9%estimated ± 3.7 pp, low confidence
Claude Opus 4.730.9%53.9%estimated ± 3.7 pp, low confidence
Claude Sonnet 538.2%53.9%estimated ± 3.7 pp, medium confidence
Command A+22.5%53.9%estimated ± 3.7 pp, low confidence
DeepSeek-R113.1%53.9%estimated ± 3.7 pp, low confidence
DeepSeek R1 Distill Qwen 32B8.4%53.9%estimated ± 3.7 pp, low confidence
DeepSeek V38.5%53.9%estimated ± 3.7 pp, low confidence
DeepSeek V3 03249.7%53.9%estimated ± 3.7 pp, low confidence
DeepSeek V3.113.7%53.9%estimated ± 3.7 pp, low confidence
DeepSeek V3.1 (Reasoning)13.5%53.9%estimated ± 3.7 pp, low confidence
DeepSeek V3.216.0%53.9%estimated ± 3.7 pp, low confidence
DeepSeek V4.1 Flash39.5%53.9%estimated ± 3.7 pp, medium confidence
Exaone 4.0 1.2B5.2%53.9%estimated ± 3.7 pp, low confidence
Exaone 4.0 32B6.3%53.9%estimated ± 3.7 pp, low confidence
Gemini 1.0 Pro5.3%53.9%estimated ± 3.7 pp, low confidence
Gemini 1.5 Pro7.9%53.9%estimated ± 3.7 pp, low confidence
Gemini 2.5 Flash9.9%53.9%estimated ± 3.7 pp, low confidence
Gemini 2.5 Pro16.1%53.9%estimated ± 3.7 pp, low confidence
Gemini 3.5 Flash-Lite22.2%53.9%estimated ± 3.7 pp, low confidence
Gemini 3 Flash17.9%53.9%estimated ± 3.7 pp, low confidence
Gemini 4 Argon52.6%59.5%estimated ± 3.7 pp, medium confidence
Gemma 3 27B4.9%53.9%estimated ± 3.7 pp, low confidence
Gemma 4 12B14.2%53.9%estimated ± 3.7 pp, low confidence
Gemma 4 26B A4B16.7%53.9%estimated ± 3.7 pp, low confidence
Gemma 4 31B14.7%53.9%estimated ± 3.7 pp, low confidence
Gemma 4 E2B7.8%53.9%estimated ± 3.7 pp, low confidence
Gemma 4 E4B8.9%53.9%estimated ± 3.7 pp, low confidence
GLM-4.5-Air11.1%53.9%estimated ± 3.7 pp, low confidence
GLM-4.614.9%53.9%estimated ± 3.7 pp, low confidence
GLM-4.722.2%53.9%estimated ± 3.7 pp, low confidence
GLM-527.9%53.9%estimated ± 3.7 pp, low confidence
GLM-5.126.1%53.9%estimated ± 3.7 pp, low confidence
GLM-5.233.7%53.9%estimated ± 3.7 pp, low confidence
GLM-5.344.8%54.0%estimated ± 3.7 pp, medium confidence
GLM-5.3-Flash41.8%53.9%estimated ± 3.7 pp, medium confidence
GLM-5-Turbo26.6%53.9%estimated ± 3.7 pp, low confidence
GLM-5V-Turbo23.5%53.9%estimated ± 3.7 pp, low confidence
GPT-4.112.7%53.9%estimated ± 3.7 pp, low confidence
GPT-4.1 mini10.2%53.9%estimated ± 3.7 pp, low confidence
GPT-4.1 nano7.8%53.9%estimated ± 3.7 pp, low confidence
GPT-4 Turbo7.0%53.9%estimated ± 3.7 pp, low confidence
GPT-4o8.4%53.9%estimated ± 3.7 pp, low confidence
GPT-4o mini6.7%53.9%estimated ± 3.7 pp, low confidence
GPT-5.1-Codex23.7%53.9%estimated ± 3.7 pp, low confidence
GPT-5.1-Codex-Max23.7%53.9%estimated ± 3.7 pp, low confidence
GPT-5.2-Codex28.5%53.9%estimated ± 3.7 pp, low confidence
GPT-5.3 Codex32.5%53.9%estimated ± 3.7 pp, low confidence
GPT-5 (high)23.0%53.9%estimated ± 3.7 pp, low confidence
GPT-5 (medium)22.9%53.9%estimated ± 3.7 pp, low confidence
GPT-OSS 120B11.6%53.9%estimated ± 3.7 pp, low confidence
GPT-OSS 20B9.0%53.9%estimated ± 3.7 pp, low confidence
Granite-4.0-350M4.8%53.9%estimated ± 3.7 pp, low confidence
Granite-4.0-H-1B5.2%53.9%estimated ± 3.7 pp, low confidence
Granite-4.0-H-350M4.8%53.9%estimated ± 3.7 pp, low confidence
Granite 4.2 30B14.8%53.9%estimated ± 3.7 pp, low confidence
Granite 4.2 3B9.1%53.9%estimated ± 3.7 pp, low confidence
Granite 4.2 8B11.1%53.9%estimated ± 3.7 pp, low confidence
Grok 422.5%53.9%estimated ± 3.7 pp, low confidence
Grok 4.1 Fast11.3%53.9%estimated ± 3.7 pp, low confidence
Grok 4.1 Fast (Reasoning)20.4%53.9%estimated ± 3.7 pp, low confidence
Grok 4.337.6%53.9%estimated ± 3.7 pp, low confidence
Grok 4 Fast (Reasoning)17.9%53.9%estimated ± 3.7 pp, low confidence
Grok Code Fast 114.1%53.9%estimated ± 3.7 pp, low confidence
Hy325.3%53.9%estimated ± 3.7 pp, low confidence
Hy3 Preview41.2%53.9%estimated ± 3.7 pp, medium confidence
Inkling25.0%53.9%estimated ± 3.7 pp, low confidence
K-Exaone14.4%53.9%estimated ± 3.7 pp, low confidence
K-EXAONE 2.019.7%53.9%estimated ± 3.7 pp, low confidence
Kimi K2.627.0%53.9%estimated ± 3.7 pp, low confidence
Kimi K212.7%53.9%estimated ± 3.7 pp, low confidence
Kimi K2.523.5%53.9%estimated ± 3.7 pp, low confidence
Kimi K2.5 (Reasoning)23.5%53.9%estimated ± 3.7 pp, low confidence
Kimi K2.7 Code25.8%53.9%estimated ± 3.7 pp, low confidence
LFM2.5-2.6B8.4%53.9%estimated ± 3.7 pp, low confidence
LFM2.5-8B-A1B7.2%53.9%estimated ± 3.7 pp, low confidence
LFM2.5-VL-1.6B-Extract4.8%53.9%estimated ± 3.7 pp, low confidence
Ling 2.6 Flash14.1%53.9%estimated ± 3.7 pp, low confidence
Ling 3.0 Flash20.1%53.9%estimated ± 3.7 pp, low confidence
Ling 3.0 Flash FP820.1%53.9%estimated ± 3.7 pp, low confidence
Ling 3.0 Flash VL24.6%53.9%estimated ± 3.7 pp, low confidence
Ling 3.0 Tiny11.1%53.9%estimated ± 3.7 pp, low confidence
Llama 3.1 405B7.3%53.9%estimated ± 3.7 pp, low confidence
Llama 4 Maverick10.0%53.9%estimated ± 3.7 pp, low confidence
Llama 4 Scout8.1%53.9%estimated ± 3.7 pp, low confidence
Mercury 2.512.3%53.9%estimated ± 3.7 pp, low confidence
MiMo-V2.5-Pro26.0%53.9%estimated ± 3.7 pp, low confidence
MiMo-V2.6-Flash37.9%53.9%estimated ± 3.7 pp, low confidence
MiMo-V2.6-Pro46.3%54.1%estimated ± 3.7 pp, medium confidence
MiMo-V2-Flash16.1%53.9%estimated ± 3.7 pp, low confidence
MiMo-V2-Omni23.9%53.9%estimated ± 3.7 pp, low confidence
MiMo-V2-Pro28.6%53.9%estimated ± 3.7 pp, low confidence
MiniCPM5-2B12.5%53.9%estimated ± 3.7 pp, low confidence
MiniMax M2.722.8%53.9%estimated ± 3.7 pp, low confidence
MiniMax M329.2%53.9%estimated ± 3.7 pp, low confidence
Mistral Large 27.6%53.9%estimated ± 3.7 pp, low confidence
Mistral Large 39.3%53.9%estimated ± 3.7 pp, low confidence
Mistral Large 438.4%53.9%estimated ± 3.7 pp, medium confidence
Mistral Medium 39.0%53.9%estimated ± 3.7 pp, low confidence
Mistral Medium 3.5 128B14.2%53.9%estimated ± 3.7 pp, low confidence
Mistral Small 411.3%53.9%estimated ± 3.7 pp, low confidence
Mistral Small 4 (Reasoning)11.3%53.9%estimated ± 3.7 pp, low confidence
Muse Glimmer 30B17.5%53.9%estimated ± 3.7 pp, low confidence
Muse Spark31.3%53.9%estimated ± 3.7 pp, low confidence
Muse Spark 1.239.6%53.9%estimated ± 3.7 pp, medium confidence
Muse Spark 1.348.1%54.4%estimated ± 3.7 pp, medium confidence
Nemotron 3.5 Lightning 30B A3B NVFP412.9%53.9%estimated ± 3.7 pp, low confidence
Nemotron 3 Nano 30B8.9%53.9%estimated ± 3.7 pp, low confidence
Nemotron 3 Nano Omni 30B A3B10.3%53.9%estimated ± 3.7 pp, low confidence
Nemotron 3 Super 100B12.8%53.9%estimated ± 3.7 pp, low confidence
Nemotron 3 Ultra22.9%53.9%estimated ± 3.7 pp, low confidence
Nemotron Ultra 253B7.5%53.9%estimated ± 3.7 pp, low confidence
North Mini Code9.9%53.9%estimated ± 3.7 pp, low confidence
Nova Pro7.0%53.9%estimated ± 3.7 pp, low confidence
o115.2%53.9%estimated ± 3.7 pp, low confidence
o1-preview11.4%53.9%estimated ± 3.7 pp, low confidence
o1-pro12.4%53.9%estimated ± 3.7 pp, low confidence
o320.2%53.9%estimated ± 3.7 pp, low confidence
o3-mini12.5%53.9%estimated ± 3.7 pp, low confidence
o3-pro21.9%53.9%estimated ± 3.7 pp, low confidence
Phi-45.9%53.9%estimated ± 3.7 pp, low confidence
Phi-4 Multimodal Instruct5.8%53.9%estimated ± 3.7 pp, low confidence
Quasar 438B26.7%53.9%estimated ± 3.7 pp, low confidence
Qwen2.5 Coder 32B Instruct6.7%53.9%estimated ± 3.7 pp, low confidence
Qwen3.5-122B-A10B15.6%53.9%estimated ± 3.7 pp, low confidence
Qwen3.5-27B22.9%53.9%estimated ± 3.7 pp, low confidence
Qwen3.5-35B-A3B19.3%53.9%estimated ± 3.7 pp, low confidence
Qwen3.5 397B21.4%53.9%estimated ± 3.7 pp, low confidence
Qwen3.5 397B (Reasoning)21.4%53.9%estimated ± 3.7 pp, low confidence
Qwen3.6-27B21.4%53.9%estimated ± 3.7 pp, low confidence
Qwen3.6-35B-A3B18.2%53.9%estimated ± 3.7 pp, low confidence
Qwen 3.6 Max (preview)28.4%53.9%estimated ± 3.7 pp, low confidence
Qwen3.6 Plus27.0%53.9%estimated ± 3.7 pp, low confidence
Qwen3.7 Max29.5%53.9%estimated ± 3.7 pp, low confidence
Qwen3.7 Plus25.2%53.9%estimated ± 3.7 pp, low confidence
Qwen3.8-27B33.7%53.9%estimated ± 3.7 pp, low confidence
Qwen3.8-Flash-Next39.8%53.9%estimated ± 3.7 pp, medium confidence
Qwen3.8 Max Preview45.4%54.0%estimated ± 3.7 pp, medium confidence
Qwen3 Max15.6%53.9%estimated ± 3.7 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct6.0%53.9%estimated ± 3.7 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking7.8%53.9%estimated ± 3.7 pp, low confidence
Sarvam 105B8.8%53.9%estimated ± 3.7 pp, low confidence
Sarvam 30B6.6%53.9%estimated ± 3.7 pp, low confidence
Solar Pro 27.0%53.9%estimated ± 3.7 pp, low confidence
Solar Pro 37.8%53.9%estimated ± 3.7 pp, low confidence
Solar Pro 428.2%53.9%estimated ± 3.7 pp, low confidence
Step 3.7 Flash19.5%53.9%estimated ± 3.7 pp, low confidence
Step 5 Preview43.7%53.9%estimated ± 3.7 pp, medium confidence
Trinity-Large-Preview10.8%53.9%estimated ± 3.7 pp, low confidence
Trinity-Large-Thinking10.8%53.9%estimated ± 3.7 pp, low confidence
Ultravox v0.6 Llama 3.3 70B7.7%53.9%estimated ± 3.7 pp, low confidence