benchgap
Calibration

JobBench → AA AutomationBench

AA AutomationBench is estimated from JobBench with a linear curve fitted on 5 models measured on both: y = 0.2804·x + 0.3954, R² = 0.53, cross-validated error 5.0 pp. It is used for 25 estimates.

Estimated modelJobBenchAA AutomationBenchSource
Atria Dawn Preview50.3%53.6%estimated ± 5.0 pp, medium confidence
Claude 4.1 Opus21.9%45.7%estimated ± 5.0 pp, low confidence
Claude 4 Sonnet18.4%44.7%estimated ± 5.0 pp, low confidence
Claude Haiku 4.516.0%44.0%estimated ± 5.0 pp, low confidence
Claude Opus 4.532.3%48.6%estimated ± 5.0 pp, low confidence
Claude Opus 4.636.7%49.8%estimated ± 5.0 pp, medium confidence
Claude Opus 4.7 (Adaptive)45.9%52.4%estimated ± 5.0 pp, medium confidence
Claude Sonnet 4.527.7%47.3%estimated ± 5.0 pp, low confidence
Claude Sonnet 4.636.9%49.9%estimated ± 5.0 pp, medium confidence
Gemini 3 Flash11.4%42.7%estimated ± 5.0 pp, low confidence
Gemini 3 Pro11.4%42.7%estimated ± 5.0 pp, low confidence
GPT-5.1-Codex26.2%46.9%estimated ± 5.0 pp, low confidence
GPT-5.234.3%49.2%estimated ± 5.0 pp, medium confidence
GPT-5.2-Codex26.0%46.8%estimated ± 5.0 pp, low confidence
GPT-5.3 Codex33.7%49.0%estimated ± 5.0 pp, medium confidence
GPT-5.438.9%50.4%estimated ± 5.0 pp, medium confidence
GPT-5.542.7%51.5%estimated ± 5.0 pp, medium confidence
GPT-5 (high)8.5%41.9%estimated ± 5.0 pp, low confidence
Hy4 preview61.7%56.8%estimated ± 5.0 pp, medium confidence
Kimi K2.58.7%42.0%estimated ± 5.0 pp, low confidence
MiMo-V2.6-Flash61.2%56.7%estimated ± 5.0 pp, medium confidence
Muse Spark 1.154.7%54.9%estimated ± 5.0 pp, medium confidence
Qwen3.5 Plus18.5%44.7%estimated ± 5.0 pp, low confidence
Qwen3.8-Flash-Next55.7%55.2%estimated ± 5.0 pp, medium confidence
Qwen3.8 Max53.4%54.5%estimated ± 5.0 pp, medium confidence