benchgap
Calibration

AutomationBench → AA-AnalystAgent

AA-AnalystAgent is estimated from AutomationBench with a Hill curve fitted on 11 models measured on both: y = 0.1324 + (0.6233 − 0.1324)·x^6.00 / (0.51875^6.00 + x^6.00), R² = 0.90, cross-validated error 7.8 pp. It is used for 14 estimates.

Estimated modelAutomationBenchAA-AnalystAgentSource
GPT-6 Astra (high)66.6%53.4%estimated ± 7.8 pp, medium confidence
Gemini 3.8 Flash (high)59.9%47.8%estimated ± 7.8 pp, medium confidence
Muse Spark 1.3 (max)57.9%45.6%estimated ± 7.8 pp, medium confidence
K2 Horizon 375B A23B37.2%19.1%estimated ± 7.8 pp, medium confidence
GPT-6 Astra (xhigh)67.2%53.8%estimated ± 7.8 pp, medium confidence
Gemini 4 Argon (high)77.5%58.3%estimated ± 7.8 pp, low confidence
GPT-6 Sol (max)61.6%49.4%estimated ± 7.8 pp, medium confidence
MiMo-V2.6-Pro58.6%46.4%estimated ± 7.8 pp, medium confidence
DeepSeek V4.1 Flash (max)68.9%54.8%estimated ± 7.8 pp, medium confidence
Mistral Large 4 Preview59.9%47.8%estimated ± 7.8 pp, medium confidence
Grok 4.7 (xhigh)65.6%52.7%estimated ± 7.8 pp, medium confidence
GPT-6 Luna (max)53.2%39.6%estimated ± 7.8 pp, medium confidence
GPT-6.1 Sol (xhigh)66.6%53.4%estimated ± 7.8 pp, medium confidence
Grok 4.6 (xhigh)67.0%53.6%estimated ± 7.8 pp, medium confidence