benchgap
Calibration

MCP Atlas → APEX-Agents

APEX-Agents is estimated from MCP Atlas with a linear curve fitted on 5 models measured on both: y = 0.8014·x + -0.3024, R² = 1.00, cross-validated error 0.6 pp. It is used for 34 estimates.

Estimated modelMCP AtlasAPEX-AgentsSource
Claude Opus 4.542.3%3.7%estimated ± 0.6 pp, low confidence
Claude Opus 4.7 (Adaptive)77.3%31.7%estimated ± 0.6 pp, medium confidence
Claude Opus 4.882.2%35.6%estimated ± 0.6 pp, medium confidence
Claude Opus 585.8%38.5%estimated ± 0.6 pp, low confidence
DeepSeek V4 Flash 073169.0%25.1%estimated ± 0.6 pp, medium confidence
DeepSeek V4 Pro 081373.6%28.7%estimated ± 0.6 pp, medium confidence
Gemini 3.5 Flash83.6%36.8%estimated ± 0.6 pp, medium confidence
GLM-531.1%0.0%estimated ± 0.6 pp, low confidence
GLM-5.171.8%27.3%estimated ± 0.6 pp, medium confidence
GLM-5.276.8%31.3%estimated ± 0.6 pp, medium confidence
GPT-5.470.6%26.3%estimated ± 0.6 pp, medium confidence
GPT-5.4 mini57.7%16.0%estimated ± 0.6 pp, low confidence
GPT-5.4 nano56.1%14.7%estimated ± 0.6 pp, low confidence
GPT-5.575.3%30.1%estimated ± 0.6 pp, medium confidence
Inkling74.1%29.1%estimated ± 0.6 pp, medium confidence
Inkling-Small79.6%33.6%estimated ± 0.6 pp, medium confidence
Kimi K2.655.9%14.6%estimated ± 0.6 pp, low confidence
Kimi K2.529.5%0.0%estimated ± 0.6 pp, low confidence
Kimi K2.7 Code76.0%30.7%estimated ± 0.6 pp, medium confidence
Ling 3.0 Flash65.5%22.3%estimated ± 0.6 pp, medium confidence
LLaDA2.2-flash46.2%6.8%estimated ± 0.6 pp, low confidence
LongCat-Flash-Lite-Sparse45.6%6.3%estimated ± 0.6 pp, low confidence
MiniMax M374.2%29.2%estimated ± 0.6 pp, medium confidence
Muse Glimmer 30B75.5%30.3%estimated ± 0.6 pp, medium confidence
Muse Spark 1.188.1%40.4%estimated ± 0.6 pp, low confidence
Ornith-1.5-35B-A3B70.2%26.0%estimated ± 0.6 pp, medium confidence
Ornith-1.5-397B80.0%33.9%estimated ± 0.6 pp, medium confidence
Ornith-1.5-9B54.2%13.2%estimated ± 0.6 pp, low confidence
Qwen3.5 397B46.1%6.7%estimated ± 0.6 pp, low confidence
Qwen3.6-35B-A3B62.8%20.1%estimated ± 0.6 pp, medium confidence
Qwen3.6 Plus48.2%8.4%estimated ± 0.6 pp, low confidence
Qwen3.7 Max76.4%31.0%estimated ± 0.6 pp, medium confidence
Qwen3.7 Plus73.2%28.4%estimated ± 0.6 pp, medium confidence
Beam78.7%32.8%estimated ± 0.6 pp, medium confidence