Calibration
AA AutomationBench → MCP Atlas
MCP Atlas is estimated from AA AutomationBench with a linear curve fitted on 5 models measured on both: y = 0.2143·x + 0.7262, R² = 0.88, cross-validated error 2.7 pp. It is used for 19 estimates.
| Estimated model | AA AutomationBench | MCP Atlas | Source |
|---|---|---|---|
| Claude Fable 5.1 | 59.4% | 85.3% | estimated ± 2.7 pp, low confidence |
| Claude Haiku 5.5 | 35.4% | 80.2% | estimated ± 2.7 pp, medium confidence |
| Claude Opus 5.5 | 69.5% | 87.5% | estimated ± 2.7 pp, low confidence |
| Claude Sonnet 5.5 | 71.8% | 88.0% | estimated ± 2.7 pp, low confidence |
| DeepSeek V4.1 Flash | 68.9% | 87.4% | estimated ± 2.7 pp, low confidence |
| Gemini 3.8 Flash | 59.9% | 85.5% | estimated ± 2.7 pp, low confidence |
| Gemini 4 Argon | 77.5% | 89.2% | estimated ± 2.7 pp, low confidence |
| GLM-5.3 | 62.2% | 85.9% | estimated ± 2.7 pp, low confidence |
| GLM-5.3-Flash | 60.4% | 85.6% | estimated ± 2.7 pp, low confidence |
| GPT-6.1 Sol | 64.9% | 86.5% | estimated ± 2.7 pp, low confidence |
| GPT-6 Astra | 68.5% | 87.3% | estimated ± 2.7 pp, low confidence |
| GPT-6 Luna | 53.2% | 84.0% | estimated ± 2.7 pp, medium confidence |
| GPT-6 Sol | 61.6% | 85.8% | estimated ± 2.7 pp, low confidence |
| Grok 4.7 | 65.6% | 86.7% | estimated ± 2.7 pp, low confidence |
| MiMo-V2.6-Pro | 58.6% | 85.2% | estimated ± 2.7 pp, low confidence |
| Mistral Large 4 | 59.9% | 85.5% | estimated ± 2.7 pp, low confidence |
| Muse Spark 1.3 | 57.9% | 85.0% | estimated ± 2.7 pp, medium confidence |
| Nemotron 3 Ultra | 3.0% | 73.3% | estimated ± 2.7 pp, low confidence |
| Qwen3.8-27B | 48.2% | 82.9% | estimated ± 2.7 pp, medium confidence |