benchgap
Calibration

MCP Atlas → AA ITBench

AA ITBench is estimated from MCP Atlas with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 0.71462·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.61, cross-validated error 3.6 pp. It is used for 26 estimates.

Estimated modelMCP AtlasAA ITBenchSource
Claude Opus 4.542.3%19.2%estimated ± 3.6 pp, low confidence
Claude Opus 4.882.2%49.9%estimated ± 3.6 pp, medium confidence
Gemini 3.5 Flash83.6%51.3%estimated ± 3.6 pp, medium confidence
GLM-531.1%13.2%estimated ± 3.6 pp, low confidence
GLM-5.171.8%40.0%estimated ± 3.6 pp, low confidence
GPT-5.470.6%39.0%estimated ± 3.6 pp, low confidence
GPT-5.4 mini57.7%29.0%estimated ± 3.6 pp, low confidence
GPT-5.4 nano56.1%27.9%estimated ± 3.6 pp, low confidence
Inkling74.1%42.1%estimated ± 3.6 pp, low confidence
Inkling-Small79.6%47.2%estimated ± 3.6 pp, medium confidence
Kimi K2.655.9%27.7%estimated ± 3.6 pp, low confidence
Kimi K2.529.5%12.4%estimated ± 3.6 pp, low confidence
Kimi K2.7 Code76.0%43.8%estimated ± 3.6 pp, medium confidence
Ling 3.0 Flash65.5%34.8%estimated ± 3.6 pp, low confidence
LLaDA2.2-flash46.2%21.5%estimated ± 3.6 pp, low confidence
LongCat-Flash-Lite-Sparse45.6%21.1%estimated ± 3.6 pp, low confidence
MiniMax M374.2%42.1%estimated ± 3.6 pp, low confidence
Muse Glimmer 30B75.5%43.3%estimated ± 3.6 pp, medium confidence
Muse Spark 1.188.1%56.3%estimated ± 3.6 pp, low confidence
Qwen3.5 397B46.1%21.4%estimated ± 3.6 pp, low confidence
Qwen3.6-35B-A3B62.8%32.7%estimated ± 3.6 pp, low confidence
Qwen3.6 Plus48.2%22.7%estimated ± 3.6 pp, low confidence
Qwen3.7 Plus73.2%41.3%estimated ± 3.6 pp, low confidence
Beam78.7%46.4%estimated ± 3.6 pp, medium confidence
Solar Open 258.2%29.3%estimated ± 3.6 pp, low confidence
Solar Pro 461.4%31.7%estimated ± 3.6 pp, low confidence