benchgap
Calibration

AA AutomationBench → MCP Atlas

MCP Atlas is estimated from AA AutomationBench with a linear curve fitted on 5 models measured on both: y = 0.2143·x + 0.7262, R² = 0.88, cross-validated error 2.7 pp. It is used for 19 estimates.

Estimated modelAA AutomationBenchMCP AtlasSource
Claude Fable 5.159.4%85.3%estimated ± 2.7 pp, low confidence
Claude Haiku 5.535.4%80.2%estimated ± 2.7 pp, medium confidence
Claude Opus 5.569.5%87.5%estimated ± 2.7 pp, low confidence
Claude Sonnet 5.571.8%88.0%estimated ± 2.7 pp, low confidence
DeepSeek V4.1 Flash68.9%87.4%estimated ± 2.7 pp, low confidence
Gemini 3.8 Flash59.9%85.5%estimated ± 2.7 pp, low confidence
Gemini 4 Argon77.5%89.2%estimated ± 2.7 pp, low confidence
GLM-5.362.2%85.9%estimated ± 2.7 pp, low confidence
GLM-5.3-Flash60.4%85.6%estimated ± 2.7 pp, low confidence
GPT-6.1 Sol64.9%86.5%estimated ± 2.7 pp, low confidence
GPT-6 Astra68.5%87.3%estimated ± 2.7 pp, low confidence
GPT-6 Luna53.2%84.0%estimated ± 2.7 pp, medium confidence
GPT-6 Sol61.6%85.8%estimated ± 2.7 pp, low confidence
Grok 4.765.6%86.7%estimated ± 2.7 pp, low confidence
MiMo-V2.6-Pro58.6%85.2%estimated ± 2.7 pp, low confidence
Mistral Large 459.9%85.5%estimated ± 2.7 pp, low confidence
Muse Spark 1.357.9%85.0%estimated ± 2.7 pp, medium confidence
Nemotron 3 Ultra3.0%73.3%estimated ± 2.7 pp, low confidence
Qwen3.8-27B48.2%82.9%estimated ± 2.7 pp, medium confidence