Calibration
ARC-AGI-1 → GPQA Diamond
GPQA Diamond is estimated from ARC-AGI-1 with a Hill curve fitted on 12 models measured on both: y = 0.8300 + (1.2000 − 0.8300)·x^6.00 / (1.10783^6.00 + x^6.00), R² = 0.69, cross-validated error 1.4 pp. It is used for 14 estimates.
| Estimated model | ARC-AGI-1 | GPQA Diamond | Source |
|---|---|---|---|
| Claude Fable 5 | 98.5% | 95.2% | estimated ± 1.4 pp, high confidence |
| Claude Fable 5.1 | 97.5% | 94.7% | estimated ± 1.4 pp, high confidence |
| Claude Opus 4.5 Thinking | 80.0% | 87.6% | estimated ± 1.4 pp, medium confidence |
| Claude Opus 4.6 (Adaptive) | 93.0% | 92.6% | estimated ± 1.4 pp, high confidence |
| Claude Sonnet 4.5 Thinking | 63.7% | 84.3% | estimated ± 1.4 pp, medium confidence |
| Gemini 3.6 Flash | 91.2% | 91.8% | estimated ± 1.4 pp, high confidence |
| Gemini 3.7 Flash | 95.5% | 93.8% | estimated ± 1.4 pp, high confidence |
| Gemini 3.8 Flash | 98.5% | 95.2% | estimated ± 1.4 pp, high confidence |
| Gemini 3 Pro | 75.0% | 86.3% | estimated ± 1.4 pp, medium confidence |
| GPT-5.1 | 72.8% | 85.8% | estimated ± 1.4 pp, medium confidence |
| GPT-5.4 Pro | 94.5% | 93.3% | estimated ± 1.4 pp, high confidence |
| GPT-5.5 Pro | 95.0% | 93.5% | estimated ± 1.4 pp, high confidence |
| Grok 4.5 | 85.7% | 89.5% | estimated ± 1.4 pp, high confidence |
| Grok 4.6 | 87.0% | 90.0% | estimated ± 1.4 pp, high confidence |