benchgap
Calibration

Vals MMLU-Pro → HealthBench Professional

HealthBench Professional is estimated from Vals MMLU-Pro with a offset logistic curve fitted on 5 models measured on both: y = 0.4983 + (0.5999 − 0.4983) / (1 + exp(−130.12·(x − 0.8576))), R² = 0.95, cross-validated error 1.3 pp. It is used for 51 estimates.

Estimated modelVals MMLU-ProHealthBench ProfessionalSource
Claude Fable 591.5%60.0%estimated ± 1.3 pp, medium confidence
Claude Fable 5.192.4%60.0%estimated ± 1.3 pp, low confidence
Claude Haiku 4.578.7%49.8%estimated ± 1.3 pp, low confidence
Claude Opus 4.789.9%59.9%estimated ± 1.3 pp, medium confidence
Claude Opus 4.889.6%59.9%estimated ± 1.3 pp, medium confidence
Claude Sonnet 4.687.3%58.8%estimated ± 1.3 pp, medium confidence
Claude Sonnet 587.5%59.0%estimated ± 1.3 pp, medium confidence
DeepSeek V4 Flash 073186.2%56.3%estimated ± 1.3 pp, medium confidence
DeepSeek V4 Pro 081387.0%58.3%estimated ± 1.3 pp, medium confidence
Gemini 3.1 Flash-Lite86.2%56.3%estimated ± 1.3 pp, medium confidence
Gemini 3.1 Pro91.0%60.0%estimated ± 1.3 pp, medium confidence
Gemini 3.5 Flash89.5%59.9%estimated ± 1.3 pp, medium confidence
Gemini 3.5 Flash-Lite85.8%55.0%estimated ± 1.3 pp, low confidence
Gemini 3.6 Flash89.3%59.9%estimated ± 1.3 pp, medium confidence
Gemini 3.7 Flash90.1%60.0%estimated ± 1.3 pp, medium confidence
Gemini 3.8 Flash90.2%60.0%estimated ± 1.3 pp, medium confidence
Gemini 3 Flash88.6%59.7%estimated ± 1.3 pp, medium confidence
GLM-4.581.2%49.9%estimated ± 1.3 pp, low confidence
GLM-4.682.2%49.9%estimated ± 1.3 pp, low confidence
GLM-4.782.7%50.0%estimated ± 1.3 pp, low confidence
GLM-5.186.9%58.1%estimated ± 1.3 pp, medium confidence
GLM-5.286.7%57.7%estimated ± 1.3 pp, medium confidence
GLM-5.386.8%57.9%estimated ± 1.3 pp, medium confidence
GLM-5.3-Flash86.1%56.0%estimated ± 1.3 pp, medium confidence
GPT-5.4 mini84.6%51.7%estimated ± 1.3 pp, low confidence
GPT-5.4 nano77.2%49.8%estimated ± 1.3 pp, low confidence
GPT-5.588.1%59.5%estimated ± 1.3 pp, medium confidence
Grok 4.2086.3%56.6%estimated ± 1.3 pp, medium confidence
Grok 4.385.8%55.0%estimated ± 1.3 pp, low confidence
Grok 4.589.2%59.9%estimated ± 1.3 pp, medium confidence
Grok 4.689.4%59.9%estimated ± 1.3 pp, medium confidence
Inkling86.3%56.6%estimated ± 1.3 pp, medium confidence
Inkling-Small85.6%54.4%estimated ± 1.3 pp, low confidence
Kimi K2.687.6%59.1%estimated ± 1.3 pp, medium confidence
Kimi K388.0%59.5%estimated ± 1.3 pp, medium confidence
Laguna M.168.8%49.8%estimated ± 1.3 pp, low confidence
Laguna XS.269.1%49.8%estimated ± 1.3 pp, low confidence
Ling 3.0 Flash82.0%49.9%estimated ± 1.3 pp, low confidence
MiMo-V2.582.9%50.1%estimated ± 1.3 pp, low confidence
MiMo-V2.5-Pro84.6%51.7%estimated ± 1.3 pp, low confidence
MiniMax M2.780.4%49.8%estimated ± 1.3 pp, low confidence
MiniMax M384.2%51.0%estimated ± 1.3 pp, low confidence
Mistral Medium 3.5 128B75.3%49.8%estimated ± 1.3 pp, low confidence
Muse Spark87.3%58.8%estimated ± 1.3 pp, medium confidence
Muse Spark 1.288.3%59.6%estimated ± 1.3 pp, medium confidence
Nemotron 3 Ultra85.8%55.0%estimated ± 1.3 pp, low confidence
Qwen3.5 Flash84.1%50.9%estimated ± 1.3 pp, low confidence
Qwen3.6 Plus87.7%59.2%estimated ± 1.3 pp, medium confidence
Qwen3.7 Max89.3%59.9%estimated ± 1.3 pp, medium confidence
Qwen3.8-27B84.3%51.2%estimated ± 1.3 pp, low confidence
Qwen3.8 Max88.6%59.7%estimated ± 1.3 pp, medium confidence