Calibration
ARC-AGI-2 → HLE-Verified
HLE-Verified is estimated from ARC-AGI-2 with a offset logistic curve fitted on 5 models measured on both: y = 0.3534 + (0.5460 − 0.3534) / (1 + exp(−200.00·(x − 0.8315))), R² = 0.99, cross-validated error 0.7 pp. It is used for 36 estimates.
| Estimated model | ARC-AGI-2 | HLE-Verified | Source |
|---|---|---|---|
| Claude Fable 5 | 89.2% | 54.6% | estimated ± 0.7 pp, medium confidence |
| Claude Fable 5.1 | 90.0% | 54.6% | estimated ± 0.7 pp, medium confidence |
| Claude Opus 4.5 Thinking | 37.6% | 35.3% | estimated ± 0.7 pp, low confidence |
| Claude Opus 4.6 (Adaptive) | 68.8% | 35.3% | estimated ± 0.7 pp, low confidence |
| Claude Opus 4.7 (Adaptive) | 75.8% | 35.3% | estimated ± 0.7 pp, low confidence |
| Claude Opus 4.8 | 72.1% | 35.3% | estimated ± 0.7 pp, low confidence |
| Claude Opus 5.5 | 91.7% | 54.6% | estimated ± 0.7 pp, medium confidence |
| Claude Sonnet 4.5 | 13.6% | 35.3% | estimated ± 0.7 pp, low confidence |
| Claude Sonnet 4.5 Thinking | 13.6% | 35.3% | estimated ± 0.7 pp, low confidence |
| Claude Sonnet 4.6 | 58.3% | 35.3% | estimated ± 0.7 pp, low confidence |
| DeepSeek V4 Flash 0731 | 61.4% | 35.3% | estimated ± 0.7 pp, low confidence |
| DeepSeek V4 Pro 0813 | 61.3% | 35.3% | estimated ± 0.7 pp, low confidence |
| dots3-note Preview | 81.4% | 35.9% | estimated ± 0.7 pp, low confidence |
| Gemini 3.1 Pro | 77.1% | 35.3% | estimated ± 0.7 pp, low confidence |
| Gemini 3.5 Flash | 72.1% | 35.3% | estimated ± 0.7 pp, low confidence |
| Gemini 3.6 Flash | 60.4% | 35.3% | estimated ± 0.7 pp, low confidence |
| Gemini 3 Pro | 31.1% | 35.3% | estimated ± 0.7 pp, low confidence |
| Gemini 3 Pro Deep Think | 45.1% | 35.3% | estimated ± 0.7 pp, low confidence |
| GPT-5.2 | 52.9% | 35.3% | estimated ± 0.7 pp, low confidence |
| GPT-5.4 | 74.0% | 35.3% | estimated ± 0.7 pp, low confidence |
| GPT-5.4 mini | 18.9% | 35.3% | estimated ± 0.7 pp, low confidence |
| GPT-5.4 nano | 5.7% | 35.3% | estimated ± 0.7 pp, low confidence |
| GPT-5.4 Pro | 83.3% | 46.4% | estimated ± 0.7 pp, low confidence |
| GPT-5.5 | 85.0% | 54.1% | estimated ± 0.7 pp, medium confidence |
| GPT-5.5 Pro | 84.2% | 52.5% | estimated ± 0.7 pp, medium confidence |
| GPT-5.6 Luna | 59.5% | 35.3% | estimated ± 0.7 pp, low confidence |
| GPT-6.1 Sol | 94.2% | 54.6% | estimated ± 0.7 pp, low confidence |
| GPT-6 Astra | 95.0% | 54.6% | estimated ± 0.7 pp, low confidence |
| GPT-6 Luna | 59.3% | 35.3% | estimated ± 0.7 pp, low confidence |
| GPT-6 Sol | 89.6% | 54.6% | estimated ± 0.7 pp, medium confidence |
| Grok 4.20 | 53.3% | 35.3% | estimated ± 0.7 pp, low confidence |
| Grok 4.5 | 52.6% | 35.3% | estimated ± 0.7 pp, low confidence |
| Grok 4.6 | 67.1% | 35.3% | estimated ± 0.7 pp, low confidence |
| Inkling-Small | 40.1% | 35.3% | estimated ± 0.7 pp, low confidence |
| Kimi K3 | 60.4% | 35.3% | estimated ± 0.7 pp, low confidence |
| Muse Spark | 42.5% | 35.3% | estimated ± 0.7 pp, low confidence |