Calibration
JobBench → Agents' Last Exam
Agents' Last Exam is estimated from JobBench with a offset logistic curve fitted on 7 models measured on both: y = 0.4889 + (0.2725 − 0.4889) / (1 + exp(−200.00·(x − 0.5798))), R² = 0.89, cross-validated error 8.7 pp. It is used for 13 estimates.
| Estimated model | JobBench | Agents' Last Exam | Source |
|---|---|---|---|
| Claude 4.1 Opus | 21.9% | 48.9% | estimated ± 8.7 pp, low confidence |
| Claude 4 Sonnet | 18.4% | 48.9% | estimated ± 8.7 pp, low confidence |
| Claude Haiku 4.5 | 16.0% | 48.9% | estimated ± 8.7 pp, low confidence |
| Claude Sonnet 4.5 | 27.7% | 48.9% | estimated ± 8.7 pp, low confidence |
| Gemini 3 Flash | 11.4% | 48.9% | estimated ± 8.7 pp, low confidence |
| Gemini 3 Pro | 11.4% | 48.9% | estimated ± 8.7 pp, low confidence |
| GPT-5.1-Codex | 26.2% | 48.9% | estimated ± 8.7 pp, low confidence |
| GPT-5.2 | 34.3% | 48.9% | estimated ± 8.7 pp, low confidence |
| GPT-5.2-Codex | 26.0% | 48.9% | estimated ± 8.7 pp, low confidence |
| GPT-5.3 Codex | 33.7% | 48.9% | estimated ± 8.7 pp, low confidence |
| GPT-5 (high) | 8.5% | 48.9% | estimated ± 8.7 pp, low confidence |
| Kimi K2.5 | 8.7% | 48.9% | estimated ± 8.7 pp, low confidence |
| Qwen3.5 Plus | 18.5% | 48.9% | estimated ± 8.7 pp, low confidence |