Calibration
ARC-AGI-1 → HealthBench (length-adjusted)
HealthBench (length-adjusted) is estimated from ARC-AGI-1 with a Michaelis–Menten curve fitted on 6 models measured on both: y = 1.5144·x / (1.57246 + x), R² = 0.36, cross-validated error 2.5 pp. It is used for 20 estimates.
| Estimated model | ARC-AGI-1 | HealthBench (length-adjusted) | Source |
|---|---|---|---|
| Claude Fable 5 | 98.5% | 58.3% | estimated ± 2.5 pp, low confidence |
| Claude Fable 5.1 | 97.5% | 58.0% | estimated ± 2.5 pp, low confidence |
| Claude Opus 4.5 Thinking | 80.0% | 51.1% | estimated ± 2.5 pp, low confidence |
| Claude Opus 4.6 (Adaptive) | 93.0% | 56.3% | estimated ± 2.5 pp, low confidence |
| Claude Sonnet 4.5 Thinking | 63.7% | 43.6% | estimated ± 2.5 pp, low confidence |
| Claude Sonnet 4.6 | 86.0% | 53.5% | estimated ± 2.5 pp, low confidence |
| DeepSeek V4 Flash 0731 | 89.0% | 54.7% | estimated ± 2.5 pp, low confidence |
| DeepSeek V4 Pro 0813 | 90.0% | 55.1% | estimated ± 2.5 pp, low confidence |
| Gemini 3.5 Flash | 92.5% | 56.1% | estimated ± 2.5 pp, low confidence |
| Gemini 3.6 Flash | 91.2% | 55.6% | estimated ± 2.5 pp, low confidence |
| Gemini 3.7 Flash | 95.5% | 57.2% | estimated ± 2.5 pp, low confidence |
| Gemini 3 Pro | 75.0% | 48.9% | estimated ± 2.5 pp, low confidence |
| GPT-5.1 | 72.8% | 47.9% | estimated ± 2.5 pp, low confidence |
| GPT-5.2 | 86.2% | 53.6% | estimated ± 2.5 pp, low confidence |
| GPT-5.4 mini | 63.7% | 43.7% | estimated ± 2.5 pp, low confidence |
| GPT-5.4 nano | 51.5% | 37.4% | estimated ± 2.5 pp, low confidence |
| GPT-5.4 Pro | 94.5% | 56.8% | estimated ± 2.5 pp, low confidence |
| GPT-5.5 Pro | 95.0% | 57.0% | estimated ± 2.5 pp, low confidence |
| Inkling-Small | 84.0% | 52.7% | estimated ± 2.5 pp, low confidence |
| Kimi K3 | 94.5% | 56.8% | estimated ± 2.5 pp, low confidence |