benchgap
Calibration

HealthBench Professional → Vals MMLU-Pro

Vals MMLU-Pro is estimated from HealthBench Professional with a Michaelis–Menten curve fitted on 5 models measured on both: y = 2.0000·x / (0.73920 + x), R² = 0.66, cross-validated error 1.5 pp. It is used for 7 estimates.

Estimated modelHealthBench ProfessionalVals MMLU-ProSource
Claude Opus 5.565.6%94.0%estimated ± 1.5 pp, low confidence
Claude Sonnet 5.569.2%96.7%estimated ± 1.5 pp, low confidence
GPT-6.1 Sol64.2%93.0%estimated ± 1.5 pp, low confidence
GPT-6 Luna60.8%90.3%estimated ± 1.5 pp, low confidence
GPT-6 Sol60.8%90.3%estimated ± 1.5 pp, low confidence
Grok 4.756.7%86.8%estimated ± 1.5 pp, medium confidence
Ling 3.1 Flash65.4%93.8%estimated ± 1.5 pp, low confidence