benchgap
Calibration

JobBench → Toolathlon-Verified

Toolathlon-Verified is estimated from JobBench with a offset logistic curve fitted on 7 models measured on both: y = 0.7318 + (1.2000 − 0.7318) / (1 + exp(−200.00·(x − 0.6331))), R² = 0.78, cross-validated error 1.3 pp. It is used for 18 estimates.

Estimated modelJobBenchToolathlon-VerifiedSource
Atria Dawn Preview50.3%73.2%estimated ± 1.3 pp, low confidence
Claude 4.1 Opus21.9%73.2%estimated ± 1.3 pp, low confidence
Claude 4 Sonnet18.4%73.2%estimated ± 1.3 pp, low confidence
Claude Haiku 4.516.0%73.2%estimated ± 1.3 pp, low confidence
Claude Opus 4.532.3%73.2%estimated ± 1.3 pp, low confidence
Claude Opus 4.636.7%73.2%estimated ± 1.3 pp, low confidence
Claude Sonnet 4.527.7%73.2%estimated ± 1.3 pp, low confidence
Gemini 3 Flash11.4%73.2%estimated ± 1.3 pp, low confidence
Gemini 3 Pro11.4%73.2%estimated ± 1.3 pp, low confidence
GPT-5.1-Codex26.2%73.2%estimated ± 1.3 pp, low confidence
GPT-5.234.3%73.2%estimated ± 1.3 pp, low confidence
GPT-5.2-Codex26.0%73.2%estimated ± 1.3 pp, low confidence
GPT-5.3 Codex33.7%73.2%estimated ± 1.3 pp, low confidence
GPT-5.438.9%73.2%estimated ± 1.3 pp, low confidence
GPT-5 (high)8.5%73.2%estimated ± 1.3 pp, low confidence
Kimi K2.58.7%73.2%estimated ± 1.3 pp, low confidence
Qwen3.5 Plus18.5%73.2%estimated ± 1.3 pp, low confidence
Qwen3.8-27B33.4%73.2%estimated ± 1.3 pp, low confidence