benchgap
Calibration

JobBench → Agents' Last Exam

Agents' Last Exam is estimated from JobBench with a offset logistic curve fitted on 7 models measured on both: y = 0.4889 + (0.2725 − 0.4889) / (1 + exp(−200.00·(x − 0.5798))), R² = 0.89, cross-validated error 8.7 pp. It is used for 13 estimates.

Estimated modelJobBenchAgents' Last ExamSource
Claude 4.1 Opus21.9%48.9%estimated ± 8.7 pp, low confidence
Claude 4 Sonnet18.4%48.9%estimated ± 8.7 pp, low confidence
Claude Haiku 4.516.0%48.9%estimated ± 8.7 pp, low confidence
Claude Sonnet 4.527.7%48.9%estimated ± 8.7 pp, low confidence
Gemini 3 Flash11.4%48.9%estimated ± 8.7 pp, low confidence
Gemini 3 Pro11.4%48.9%estimated ± 8.7 pp, low confidence
GPT-5.1-Codex26.2%48.9%estimated ± 8.7 pp, low confidence
GPT-5.234.3%48.9%estimated ± 8.7 pp, low confidence
GPT-5.2-Codex26.0%48.9%estimated ± 8.7 pp, low confidence
GPT-5.3 Codex33.7%48.9%estimated ± 8.7 pp, low confidence
GPT-5 (high)8.5%48.9%estimated ± 8.7 pp, low confidence
Kimi K2.58.7%48.9%estimated ± 8.7 pp, low confidence
Qwen3.5 Plus18.5%48.9%estimated ± 8.7 pp, low confidence