benchgap
Calibration

ARC-AGI-2 → HealthBench Professional

HealthBench Professional is estimated from ARC-AGI-2 with a offset logistic curve fitted on 10 models measured on both: y = 0.5510 + (0.6587 − 0.5510) / (1 + exp(−32.69·(x − 0.8938))), R² = 0.57, cross-validated error 4.6 pp. It is used for 10 estimates.

Estimated modelARC-AGI-2HealthBench ProfessionalSource
Claude Opus 4.5 Thinking37.6%55.1%estimated ± 4.6 pp, medium confidence
Claude Opus 4.6 (Adaptive)68.8%55.1%estimated ± 4.6 pp, high confidence
Claude Sonnet 4.513.6%55.1%estimated ± 4.6 pp, medium confidence
Claude Sonnet 4.5 Thinking13.6%55.1%estimated ± 4.6 pp, medium confidence
dots3-note Preview81.4%55.8%estimated ± 4.6 pp, high confidence
Gemini 3 Pro31.1%55.1%estimated ± 4.6 pp, medium confidence
Gemini 3 Pro Deep Think45.1%55.1%estimated ± 4.6 pp, medium confidence
GPT-5.252.9%55.1%estimated ± 4.6 pp, medium confidence
GPT-5.4 Pro83.3%56.4%estimated ± 4.6 pp, high confidence
GPT-5.5 Pro84.2%56.8%estimated ± 4.6 pp, high confidence