Calibration
JobBench → MCP Atlas
MCP Atlas is estimated from JobBench with a linear curve fitted on 9 models measured on both: y = 1.1904·x + 0.1824, R² = 0.88, cross-validated error 8.2 pp. It is used for 15 estimates.
| Estimated model | JobBench | MCP Atlas | Source |
|---|---|---|---|
| Atria Dawn Preview | 50.3% | 78.1% | estimated ± 8.2 pp, medium confidence |
| Claude 4.1 Opus | 21.9% | 44.3% | estimated ± 8.2 pp, medium confidence |
| Claude 4 Sonnet | 18.4% | 40.1% | estimated ± 8.2 pp, medium confidence |
| Claude Haiku 4.5 | 16.0% | 37.3% | estimated ± 8.2 pp, medium confidence |
| Claude Opus 4.6 | 36.7% | 61.9% | estimated ± 8.2 pp, medium confidence |
| Claude Sonnet 4.5 | 27.7% | 51.2% | estimated ± 8.2 pp, medium confidence |
| Claude Sonnet 4.6 | 36.9% | 62.2% | estimated ± 8.2 pp, medium confidence |
| Gemini 3 Flash | 11.4% | 31.8% | estimated ± 8.2 pp, medium confidence |
| Gemini 3 Pro | 11.4% | 31.8% | estimated ± 8.2 pp, medium confidence |
| GPT-5.1-Codex | 26.2% | 49.4% | estimated ± 8.2 pp, medium confidence |
| GPT-5.2 | 34.3% | 59.1% | estimated ± 8.2 pp, medium confidence |
| GPT-5.2-Codex | 26.0% | 49.2% | estimated ± 8.2 pp, medium confidence |
| GPT-5.3 Codex | 33.7% | 58.4% | estimated ± 8.2 pp, medium confidence |
| GPT-5 (high) | 8.5% | 28.4% | estimated ± 8.2 pp, low confidence |
| Qwen3.5 Plus | 18.5% | 40.3% | estimated ± 8.2 pp, medium confidence |