Calibration
AutomationBench → AA-AnalystAgent
AA-AnalystAgent is estimated from AutomationBench with a Hill curve fitted on 11 models measured on both: y = 0.1324 + (0.6233 − 0.1324)·x^6.00 / (0.51875^6.00 + x^6.00), R² = 0.90, cross-validated error 7.8 pp. It is used for 14 estimates.
| Estimated model | AutomationBench | AA-AnalystAgent | Source |
|---|---|---|---|
| GPT-6 Astra (high) | 66.6% | 53.4% | estimated ± 7.8 pp, medium confidence |
| Gemini 3.8 Flash (high) | 59.9% | 47.8% | estimated ± 7.8 pp, medium confidence |
| Muse Spark 1.3 (max) | 57.9% | 45.6% | estimated ± 7.8 pp, medium confidence |
| K2 Horizon 375B A23B | 37.2% | 19.1% | estimated ± 7.8 pp, medium confidence |
| GPT-6 Astra (xhigh) | 67.2% | 53.8% | estimated ± 7.8 pp, medium confidence |
| Gemini 4 Argon (high) | 77.5% | 58.3% | estimated ± 7.8 pp, low confidence |
| GPT-6 Sol (max) | 61.6% | 49.4% | estimated ± 7.8 pp, medium confidence |
| MiMo-V2.6-Pro | 58.6% | 46.4% | estimated ± 7.8 pp, medium confidence |
| DeepSeek V4.1 Flash (max) | 68.9% | 54.8% | estimated ± 7.8 pp, medium confidence |
| Mistral Large 4 Preview | 59.9% | 47.8% | estimated ± 7.8 pp, medium confidence |
| Grok 4.7 (xhigh) | 65.6% | 52.7% | estimated ± 7.8 pp, medium confidence |
| GPT-6 Luna (max) | 53.2% | 39.6% | estimated ± 7.8 pp, medium confidence |
| GPT-6.1 Sol (xhigh) | 66.6% | 53.4% | estimated ± 7.8 pp, medium confidence |
| Grok 4.6 (xhigh) | 67.0% | 53.6% | estimated ± 7.8 pp, medium confidence |