Calibration
GDP.pdf → Agents' Last Exam
Agents' Last Exam is estimated from GDP.pdf with a Michaelis–Menten curve fitted on 9 models measured on both: y = 2.0000·x / (0.76892 + x), R² = 0.77, cross-validated error 6.3 pp. It is used for 15 estimates.
| Estimated model | GDP.pdf | Agents' Last Exam | Source |
|---|---|---|---|
| Claude Fable 5.1 | 26.2% | 50.8% | estimated ± 6.3 pp, medium confidence |
| Claude Haiku 5.5 | 20.8% | 42.6% | estimated ± 6.3 pp, medium confidence |
| Claude Opus 5.5 | 26.2% | 50.8% | estimated ± 6.3 pp, medium confidence |
| Claude Sonnet 5.5 | 25.8% | 50.2% | estimated ± 6.3 pp, medium confidence |
| Gemini 3.8 Flash | 21.0% | 42.9% | estimated ± 6.3 pp, medium confidence |
| GPT-6.1 Sol | 31.0% | 57.5% | estimated ± 6.3 pp, medium confidence |
| GPT-6 Luna | 22.8% | 45.7% | estimated ± 6.3 pp, medium confidence |
| Grok 4.7 | 20.0% | 41.3% | estimated ± 6.3 pp, medium confidence |
| Inkling | 12.8% | 28.5% | estimated ± 6.3 pp, medium confidence |
| Kimi K3 | 22.0% | 44.5% | estimated ± 6.3 pp, medium confidence |
| MiniMax M3 | 9.8% | 22.6% | estimated ± 6.3 pp, low confidence |
| Mistral Large 4 | 18.6% | 39.0% | estimated ± 6.3 pp, medium confidence |
| Muse Glimmer 30B | 10.0% | 23.0% | estimated ± 6.3 pp, low confidence |
| Muse Spark 1.3 | 26.6% | 51.4% | estimated ± 6.3 pp, medium confidence |
| Nemotron 3 Ultra | 5.0% | 12.2% | estimated ± 6.3 pp, low confidence |