Calibration
ARC-AGI-1 → LABBench2
LABBench2 is estimated from ARC-AGI-1 with a Hill curve fitted on 5 models measured on both: y = 0.0639 + (1.2000 − 0.0639)·x^6.00 / (0.85725^6.00 + x^6.00), R² = 0.77, cross-validated error 1.3 pp. It is used for 30 estimates.
| Estimated model | ARC-AGI-1 | LABBench2 | Source |
|---|---|---|---|
| Claude Fable 5 | 98.5% | 85.6% | estimated ± 1.3 pp, medium confidence |
| Claude Fable 5.1 | 97.5% | 84.1% | estimated ± 1.3 pp, medium confidence |
| Claude Opus 4.5 Thinking | 80.0% | 51.6% | estimated ± 1.3 pp, low confidence |
| Claude Opus 4.6 (Adaptive) | 93.0% | 76.8% | estimated ± 1.3 pp, low confidence |
| Claude Opus 4.8 | 92.5% | 75.9% | estimated ± 1.3 pp, low confidence |
| Claude Opus 5.5 | 97.5% | 84.1% | estimated ± 1.3 pp, medium confidence |
| Claude Sonnet 4.5 Thinking | 63.7% | 22.7% | estimated ± 1.3 pp, low confidence |
| Claude Sonnet 4.6 | 86.0% | 63.7% | estimated ± 1.3 pp, low confidence |
| DeepSeek V4 Flash 0731 | 89.0% | 69.6% | estimated ± 1.3 pp, low confidence |
| DeepSeek V4 Pro 0813 | 90.0% | 71.4% | estimated ± 1.3 pp, low confidence |
| Gemini 3.5 Flash | 92.5% | 75.9% | estimated ± 1.3 pp, low confidence |
| Gemini 3.6 Flash | 91.2% | 73.6% | estimated ± 1.3 pp, low confidence |
| Gemini 3 Pro | 75.0% | 41.6% | estimated ± 1.3 pp, low confidence |
| GPT-5.1 | 72.8% | 37.4% | estimated ± 1.3 pp, low confidence |
| GPT-5.2 | 86.2% | 64.1% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 | 93.7% | 78.0% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 mini | 63.7% | 22.8% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 nano | 51.5% | 11.5% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 Pro | 94.5% | 79.3% | estimated ± 1.3 pp, low confidence |
| GPT-5.5 | 95.0% | 80.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.5 Pro | 95.0% | 80.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.6 Luna | 88.0% | 67.6% | estimated ± 1.3 pp, low confidence |
| GPT-6.1 Sol | 96.5% | 82.6% | estimated ± 1.3 pp, medium confidence |
| GPT-6 Astra | 98.5% | 85.6% | estimated ± 1.3 pp, medium confidence |
| GPT-6 Luna | 86.7% | 65.1% | estimated ± 1.3 pp, low confidence |
| GPT-6 Sol | 95.5% | 81.0% | estimated ± 1.3 pp, medium confidence |
| Grok 4.5 | 85.7% | 63.1% | estimated ± 1.3 pp, low confidence |
| Grok 4.6 | 87.0% | 65.7% | estimated ± 1.3 pp, low confidence |
| Inkling-Small | 84.0% | 59.7% | estimated ± 1.3 pp, low confidence |
| Kimi K3 | 94.5% | 79.3% | estimated ± 1.3 pp, low confidence |