benchgap
Calibration

AA-AnalystAgent → AutomationBench

AutomationBench is estimated from AA-AnalystAgent with a inverse Michaelis–Menten curve fitted on 11 models measured on both: y = 1.83492·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.83, cross-validated error 11.2 pp. It is used for 5 estimates.

Estimated modelAA-AnalystAgentAutomationBenchSource
GPT-5.6 Sol (max)47.5%57.2%estimated ± 11.2 pp, low confidence
Claude Fable 5 (with fallback)48.8%59.2%estimated ± 11.2 pp, low confidence
Gemini 3.7 Flash (high)60.0%78.6%estimated ± 11.2 pp, low confidence
Claude Opus 5 (max)53.8%67.5%estimated ± 11.2 pp, low confidence
GPT-5.5 (xhigh)50.0%61.2%estimated ± 11.2 pp, low confidence