benchgap
Calibration

GDP.pdf → Agents' Last Exam

Agents' Last Exam is estimated from GDP.pdf with a Michaelis–Menten curve fitted on 9 models measured on both: y = 2.0000·x / (0.76892 + x), R² = 0.77, cross-validated error 6.3 pp. It is used for 15 estimates.

Estimated modelGDP.pdfAgents' Last ExamSource
Claude Fable 5.126.2%50.8%estimated ± 6.3 pp, medium confidence
Claude Haiku 5.520.8%42.6%estimated ± 6.3 pp, medium confidence
Claude Opus 5.526.2%50.8%estimated ± 6.3 pp, medium confidence
Claude Sonnet 5.525.8%50.2%estimated ± 6.3 pp, medium confidence
Gemini 3.8 Flash21.0%42.9%estimated ± 6.3 pp, medium confidence
GPT-6.1 Sol31.0%57.5%estimated ± 6.3 pp, medium confidence
GPT-6 Luna22.8%45.7%estimated ± 6.3 pp, medium confidence
Grok 4.720.0%41.3%estimated ± 6.3 pp, medium confidence
Inkling12.8%28.5%estimated ± 6.3 pp, medium confidence
Kimi K322.0%44.5%estimated ± 6.3 pp, medium confidence
MiniMax M39.8%22.6%estimated ± 6.3 pp, low confidence
Mistral Large 418.6%39.0%estimated ± 6.3 pp, medium confidence
Muse Glimmer 30B10.0%23.0%estimated ± 6.3 pp, low confidence
Muse Spark 1.326.6%51.4%estimated ± 6.3 pp, medium confidence
Nemotron 3 Ultra5.0%12.2%estimated ± 6.3 pp, low confidence