Calibration
Humanity's Last Exam → GPQA Diamond
GPQA Diamond is estimated from Humanity's Last Exam with a Michaelis–Menten curve fitted on 14 models measured on both: y = 1.0536·x / (0.05994 + x), R² = 0.90, cross-validated error 1.3 pp. It is used for 16 estimates.
| Estimated model | Humanity's Last Exam | GPQA Diamond | Source |
|---|---|---|---|
| Claude Fable 5.1 (xhigh with fallback) | 58.7% | 95.6% | estimated ± 1.3 pp, high confidence |
| Claude Fable 5.1 (high with fallback) | 55.9% | 95.2% | estimated ± 1.3 pp, high confidence |
| Claude Sonnet 5.5 (max with fallback) | 55.0% | 95.0% | estimated ± 1.3 pp, high confidence |
| Claude Opus 5.5 (max with fallback) | 61.4% | 96.0% | estimated ± 1.3 pp, medium confidence |
| Claude Opus 5.5 (xhigh with fallback) | 57.5% | 95.4% | estimated ± 1.3 pp, high confidence |
| Gemini 4 Argon (high) | 57.1% | 95.3% | estimated ± 1.3 pp, high confidence |
| Claude Opus 5.5 (high with fallback) | 55.6% | 95.1% | estimated ± 1.3 pp, high confidence |
| GPT-6.1 Sol (max) | 52.9% | 94.6% | estimated ± 1.3 pp, high confidence |
| GPT-6 Sol (max) | 47.9% | 93.6% | estimated ± 1.3 pp, high confidence |
| MiMo-V2.6-Pro | 49.4% | 94.0% | estimated ± 1.3 pp, high confidence |
| Step 5 Preview | 46.5% | 93.3% | estimated ± 1.3 pp, high confidence |
| DeepSeek V4.1 Flash (max) | 39.2% | 91.4% | estimated ± 1.3 pp, high confidence |
| Mistral Large 4 Preview | 35.0% | 90.0% | estimated ± 1.3 pp, high confidence |
| Grok 4.7 (xhigh) | 43.1% | 92.5% | estimated ± 1.3 pp, high confidence |
| GPT-6 Luna (max) | 38.5% | 91.2% | estimated ± 1.3 pp, high confidence |
| Claude Fable 5 (with fallback) | 55.5% | 95.1% | estimated ± 1.3 pp, high confidence |