benchgap
Calibration

MCP Atlas → DeepPlanning

DeepPlanning is estimated from MCP Atlas with a Michaelis–Menten + offset curve fitted on 7 models measured on both: y = 0.0000 + 2.0000·x / (2.39640 + x), R² = 0.58, cross-validated error 13.1 pp. It is used for 29 estimates.

Estimated modelMCP AtlasDeepPlanningSource
Claude Opus 4.7 (Adaptive)77.3%48.8%estimated ± 13.1 pp, low confidence
Claude Opus 4.882.2%51.1%estimated ± 13.1 pp, low confidence
Claude Opus 585.8%52.7%estimated ± 13.1 pp, low confidence
DeepSeek V4 Flash 073169.0%44.7%estimated ± 13.1 pp, low confidence
DeepSeek V4 Pro 081373.6%47.0%estimated ± 13.1 pp, low confidence
Gemini 3.5 Flash83.6%51.7%estimated ± 13.1 pp, low confidence
GLM-5.276.8%48.5%estimated ± 13.1 pp, low confidence
GPT-5.470.6%45.5%estimated ± 13.1 pp, low confidence
GPT-5.4 mini57.7%38.8%estimated ± 13.1 pp, low confidence
GPT-5.4 nano56.1%37.9%estimated ± 13.1 pp, low confidence
GPT-5.575.3%47.8%estimated ± 13.1 pp, low confidence
Hy4 preview83.7%51.8%estimated ± 13.1 pp, low confidence
Inkling74.1%47.2%estimated ± 13.1 pp, low confidence
Inkling-Small79.6%49.9%estimated ± 13.1 pp, low confidence
Kimi K2.655.9%37.8%estimated ± 13.1 pp, low confidence
Kimi K2.7 Code76.0%48.2%estimated ± 13.1 pp, low confidence
Kimi K384.2%52.0%estimated ± 13.1 pp, low confidence
Ling 3.0 Flash65.5%42.9%estimated ± 13.1 pp, low confidence
LLaDA2.2-flash46.2%32.3%estimated ± 13.1 pp, low confidence
MiniMax M374.2%47.3%estimated ± 13.1 pp, low confidence
Muse Glimmer 30B75.5%47.9%estimated ± 13.1 pp, low confidence
Muse Spark 1.188.1%53.8%estimated ± 13.1 pp, low confidence
Ornith-1.5-35B-A3B70.2%45.3%estimated ± 13.1 pp, low confidence
Ornith-1.5-397B80.0%50.1%estimated ± 13.1 pp, low confidence
Ornith-1.5-9B54.2%36.9%estimated ± 13.1 pp, low confidence
Beam78.7%49.4%estimated ± 13.1 pp, low confidence
Solar Open 258.2%39.1%estimated ± 13.1 pp, low confidence
Solar Pro 461.4%40.8%estimated ± 13.1 pp, low confidence
Step 5 Preview85.6%52.6%estimated ± 13.1 pp, low confidence