benchgap
Calibration

GPQA → HealthBench Professional

HealthBench Professional is estimated from GPQA with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (1.2000 − 0.0000)·x^6.00 / (0.95106^6.00 + x^6.00), R² = 0.53, cross-validated error 4.8 pp. It is used for 54 estimates.

Estimated modelGPQAHealthBench ProfessionalSource
Claude 3.5 Sonnet59.4%6.7%estimated ± 4.8 pp, low confidence
Claude Mythos 594.1%58.1%estimated ± 4.8 pp, medium confidence
Claude Opus 4.587.0%44.3%estimated ± 4.8 pp, low confidence
Claude Opus 4.691.3%52.7%estimated ± 4.8 pp, low confidence
DeepSeek V359.1%6.5%estimated ± 4.8 pp, low confidence
DeepSeek V4.1 Flash90.9%51.9%estimated ± 4.8 pp, low confidence
Gemini 2.5 Pro83.0%36.8%estimated ± 4.8 pp, low confidence
Gemma 4 12B78.8%29.3%estimated ± 4.8 pp, low confidence
Gemma 4 31B84.3%39.2%estimated ± 4.8 pp, low confidence
Gemma 4 E2B43.4%1.1%estimated ± 4.8 pp, low confidence
Gemma 4 E4B58.6%6.2%estimated ± 4.8 pp, low confidence
GLM-586.0%42.4%estimated ± 4.8 pp, low confidence
GPT-4.166.3%12.4%estimated ± 4.8 pp, low confidence
GPT-4.1 mini64.2%10.4%estimated ± 4.8 pp, low confidence
GPT-4.1 nano50.3%2.6%estimated ± 4.8 pp, low confidence
Granite 4.2 30B66.4%12.5%estimated ± 4.8 pp, low confidence
Granite 4.2 3B54.8%4.2%estimated ± 4.8 pp, low confidence
Granite 4.2 8B64.1%10.3%estimated ± 4.8 pp, low confidence
Hy3 Preview87.2%44.7%estimated ± 4.8 pp, low confidence
Hy4 preview92.3%54.6%estimated ± 4.8 pp, medium confidence
Interfaze Beta89.9%50.0%estimated ± 4.8 pp, low confidence
Kimi K2.587.6%45.5%estimated ± 4.8 pp, low confidence
Kimi K2.5 (Reasoning)87.6%45.5%estimated ± 4.8 pp, low confidence
LFM2.5-230M25.4%0.0%estimated ± 4.8 pp, low confidence
LFM2.5-VL-450M25.7%0.0%estimated ± 4.8 pp, low confidence
Ling 2.6 Flash59.0%6.5%estimated ± 4.8 pp, low confidence
Ling 3.0 Flash FP884.0%38.6%estimated ± 4.8 pp, low confidence
MAI-Thinking-184.2%39.0%estimated ± 4.8 pp, low confidence
Mellum2-12B-A2.5B-Instruct40.9%0.8%estimated ± 4.8 pp, low confidence
Mellum2-12B-A2.5B-Thinking57.6%5.6%estimated ± 4.8 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP475.6%24.1%estimated ± 4.8 pp, low confidence
Nemotron 3 Nano Omni 30B A3B72.2%19.3%estimated ± 4.8 pp, low confidence
o175.7%24.3%estimated ± 4.8 pp, low confidence
o1-pro79.0%29.7%estimated ± 4.8 pp, low confidence
o3-mini77.2%26.7%estimated ± 4.8 pp, low confidence
Ornith-1.5-35B-A3B89.2%48.6%estimated ± 4.8 pp, low confidence
Ornith-1.5-397B92.8%55.6%estimated ± 4.8 pp, medium confidence
Ornith-1.5-9B86.4%43.2%estimated ± 4.8 pp, low confidence
Qwen3 235B 250777.5%27.2%estimated ± 4.8 pp, low confidence
Qwen3.5-122B-A10B86.6%43.6%estimated ± 4.8 pp, low confidence
Qwen3.5-27B85.5%41.5%estimated ± 4.8 pp, low confidence
Qwen3.5-35B-A3B84.2%39.0%estimated ± 4.8 pp, low confidence
Qwen3.5 397B88.4%47.0%estimated ± 4.8 pp, low confidence
Qwen3.6-27B87.8%45.9%estimated ± 4.8 pp, low confidence
Qwen3.6-35B-A3B86.0%42.4%estimated ± 4.8 pp, low confidence
Qwen3.7 Plus90.3%50.7%estimated ± 4.8 pp, low confidence
Qwen3.8-Flash-Next91.7%53.5%estimated ± 4.8 pp, low confidence
Qwen3.8-Omni-Flash91.0%52.1%estimated ± 4.8 pp, low confidence
Sakana Fugu95.5%60.7%estimated ± 4.8 pp, medium confidence
Sakana Fugu-Ultra95.5%60.7%estimated ± 4.8 pp, medium confidence
Soofi S 30B-A3B43.4%1.1%estimated ± 4.8 pp, low confidence
Ternary Bonsai 2 27B85.8%42.0%estimated ± 4.8 pp, low confidence
ZAYA1-74B-Preview57.3%5.5%estimated ± 4.8 pp, low confidence
ZAYA1-8B71.0%17.7%estimated ± 4.8 pp, low confidence