Calibration
JobBench → Toolathlon-Verified
Toolathlon-Verified is estimated from JobBench with a offset logistic curve fitted on 7 models measured on both: y = 0.7318 + (1.2000 − 0.7318) / (1 + exp(−200.00·(x − 0.6331))), R² = 0.78, cross-validated error 1.3 pp. It is used for 18 estimates.
| Estimated model | JobBench | Toolathlon-Verified | Source |
|---|---|---|---|
| Atria Dawn Preview | 50.3% | 73.2% | estimated ± 1.3 pp, low confidence |
| Claude 4.1 Opus | 21.9% | 73.2% | estimated ± 1.3 pp, low confidence |
| Claude 4 Sonnet | 18.4% | 73.2% | estimated ± 1.3 pp, low confidence |
| Claude Haiku 4.5 | 16.0% | 73.2% | estimated ± 1.3 pp, low confidence |
| Claude Opus 4.5 | 32.3% | 73.2% | estimated ± 1.3 pp, low confidence |
| Claude Opus 4.6 | 36.7% | 73.2% | estimated ± 1.3 pp, low confidence |
| Claude Sonnet 4.5 | 27.7% | 73.2% | estimated ± 1.3 pp, low confidence |
| Gemini 3 Flash | 11.4% | 73.2% | estimated ± 1.3 pp, low confidence |
| Gemini 3 Pro | 11.4% | 73.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.1-Codex | 26.2% | 73.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.2 | 34.3% | 73.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.2-Codex | 26.0% | 73.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.3 Codex | 33.7% | 73.2% | estimated ± 1.3 pp, low confidence |
| GPT-5.4 | 38.9% | 73.2% | estimated ± 1.3 pp, low confidence |
| GPT-5 (high) | 8.5% | 73.2% | estimated ± 1.3 pp, low confidence |
| Kimi K2.5 | 8.7% | 73.2% | estimated ± 1.3 pp, low confidence |
| Qwen3.5 Plus | 18.5% | 73.2% | estimated ± 1.3 pp, low confidence |
| Qwen3.8-27B | 33.4% | 73.2% | estimated ± 1.3 pp, low confidence |