Calibration
ARC-AGI-2 → HealthBench Professional
HealthBench Professional is estimated from ARC-AGI-2 with a offset logistic curve fitted on 10 models measured on both: y = 0.5510 + (0.6587 − 0.5510) / (1 + exp(−32.69·(x − 0.8938))), R² = 0.57, cross-validated error 4.6 pp. It is used for 10 estimates.
| Estimated model | ARC-AGI-2 | HealthBench Professional | Source |
|---|---|---|---|
| Claude Opus 4.5 Thinking | 37.6% | 55.1% | estimated ± 4.6 pp, medium confidence |
| Claude Opus 4.6 (Adaptive) | 68.8% | 55.1% | estimated ± 4.6 pp, high confidence |
| Claude Sonnet 4.5 | 13.6% | 55.1% | estimated ± 4.6 pp, medium confidence |
| Claude Sonnet 4.5 Thinking | 13.6% | 55.1% | estimated ± 4.6 pp, medium confidence |
| dots3-note Preview | 81.4% | 55.8% | estimated ± 4.6 pp, high confidence |
| Gemini 3 Pro | 31.1% | 55.1% | estimated ± 4.6 pp, medium confidence |
| Gemini 3 Pro Deep Think | 45.1% | 55.1% | estimated ± 4.6 pp, medium confidence |
| GPT-5.2 | 52.9% | 55.1% | estimated ± 4.6 pp, medium confidence |
| GPT-5.4 Pro | 83.3% | 56.4% | estimated ± 4.6 pp, high confidence |
| GPT-5.5 Pro | 84.2% | 56.8% | estimated ± 4.6 pp, high confidence |