Calibration
JobBench → AA AutomationBench
AA AutomationBench is estimated from JobBench with a linear curve fitted on 5 models measured on both: y = 0.2804·x + 0.3954, R² = 0.53, cross-validated error 5.0 pp. It is used for 25 estimates.
| Estimated model | JobBench | AA AutomationBench | Source |
|---|---|---|---|
| Atria Dawn Preview | 50.3% | 53.6% | estimated ± 5.0 pp, medium confidence |
| Claude 4.1 Opus | 21.9% | 45.7% | estimated ± 5.0 pp, low confidence |
| Claude 4 Sonnet | 18.4% | 44.7% | estimated ± 5.0 pp, low confidence |
| Claude Haiku 4.5 | 16.0% | 44.0% | estimated ± 5.0 pp, low confidence |
| Claude Opus 4.5 | 32.3% | 48.6% | estimated ± 5.0 pp, low confidence |
| Claude Opus 4.6 | 36.7% | 49.8% | estimated ± 5.0 pp, medium confidence |
| Claude Opus 4.7 (Adaptive) | 45.9% | 52.4% | estimated ± 5.0 pp, medium confidence |
| Claude Sonnet 4.5 | 27.7% | 47.3% | estimated ± 5.0 pp, low confidence |
| Claude Sonnet 4.6 | 36.9% | 49.9% | estimated ± 5.0 pp, medium confidence |
| Gemini 3 Flash | 11.4% | 42.7% | estimated ± 5.0 pp, low confidence |
| Gemini 3 Pro | 11.4% | 42.7% | estimated ± 5.0 pp, low confidence |
| GPT-5.1-Codex | 26.2% | 46.9% | estimated ± 5.0 pp, low confidence |
| GPT-5.2 | 34.3% | 49.2% | estimated ± 5.0 pp, medium confidence |
| GPT-5.2-Codex | 26.0% | 46.8% | estimated ± 5.0 pp, low confidence |
| GPT-5.3 Codex | 33.7% | 49.0% | estimated ± 5.0 pp, medium confidence |
| GPT-5.4 | 38.9% | 50.4% | estimated ± 5.0 pp, medium confidence |
| GPT-5.5 | 42.7% | 51.5% | estimated ± 5.0 pp, medium confidence |
| GPT-5 (high) | 8.5% | 41.9% | estimated ± 5.0 pp, low confidence |
| Hy4 preview | 61.7% | 56.8% | estimated ± 5.0 pp, medium confidence |
| Kimi K2.5 | 8.7% | 42.0% | estimated ± 5.0 pp, low confidence |
| MiMo-V2.6-Flash | 61.2% | 56.7% | estimated ± 5.0 pp, medium confidence |
| Muse Spark 1.1 | 54.7% | 54.9% | estimated ± 5.0 pp, medium confidence |
| Qwen3.5 Plus | 18.5% | 44.7% | estimated ± 5.0 pp, low confidence |
| Qwen3.8-Flash-Next | 55.7% | 55.2% | estimated ± 5.0 pp, medium confidence |
| Qwen3.8 Max | 53.4% | 54.5% | estimated ± 5.0 pp, medium confidence |