Calibration
ARC-AGI-3 → HealthBench (length-adjusted)
HealthBench (length-adjusted) is estimated from ARC-AGI-3 with a Hill curve fitted on 5 models measured on both: y = 0.5385 + (0.5842 − 0.5385)·x^6.00 / (0.22103^6.00 + x^6.00), R² = 0.96, cross-validated error 0.9 pp. It is used for 12 estimates.
| Estimated model | ARC-AGI-3 | HealthBench (length-adjusted) | Source |
|---|---|---|---|
| Claude Opus 4.7 (Adaptive) | 0.2% | 53.8% | estimated ± 0.9 pp, medium confidence |
| Claude Opus 4.8 | 1.5% | 53.8% | estimated ± 0.9 pp, medium confidence |
| Gemini 3.1 Pro | 0.4% | 53.8% | estimated ± 0.9 pp, medium confidence |
| Gemini 3.8 Flash | 10.4% | 53.9% | estimated ± 0.9 pp, medium confidence |
| GPT-5.4 | 0.2% | 53.8% | estimated ± 0.9 pp, medium confidence |
| GPT-5.5 | 0.4% | 53.8% | estimated ± 0.9 pp, medium confidence |
| GPT-5.6 Luna | 0.2% | 53.8% | estimated ± 0.9 pp, medium confidence |
| GPT-5.6 Sol | 7.8% | 53.9% | estimated ± 0.9 pp, medium confidence |
| GPT-5.6 Terra | 0.8% | 53.8% | estimated ± 0.9 pp, medium confidence |
| Grok 4.20 | 0.1% | 53.8% | estimated ± 0.9 pp, low confidence |
| Grok 4.5 | 0.3% | 53.8% | estimated ± 0.9 pp, medium confidence |
| Grok 4.6 | 2.1% | 53.8% | estimated ± 0.9 pp, medium confidence |