benchgap
Calibration

MCP Atlas → ResearchClawBench

ResearchClawBench is estimated from MCP Atlas with a linear curve fitted on 12 models measured on both: y = 0.0958·x + 0.1144, R² = 0.49, cross-validated error 1.8 pp. It is used for 19 estimates.

Estimated modelMCP AtlasResearchClawBenchSource
Claude Opus 4.7 (Adaptive)77.3%18.8%estimated ± 1.8 pp, medium confidence
Claude Opus 585.8%19.7%estimated ± 1.8 pp, low confidence
DeepSeek V4 Flash 073169.0%18.0%estimated ± 1.8 pp, medium confidence
DeepSeek V4 Pro 081373.6%18.5%estimated ± 1.8 pp, medium confidence
GPT-5.4 mini57.7%17.0%estimated ± 1.8 pp, medium confidence
GPT-5.4 nano56.1%16.8%estimated ± 1.8 pp, medium confidence
Hy4 preview83.7%19.5%estimated ± 1.8 pp, low confidence
Inkling74.1%18.5%estimated ± 1.8 pp, medium confidence
Inkling-Small79.6%19.1%estimated ± 1.8 pp, medium confidence
Kimi K2.7 Code76.0%18.7%estimated ± 1.8 pp, medium confidence
Kimi K384.2%19.5%estimated ± 1.8 pp, low confidence
Ling 3.0 Flash65.5%17.7%estimated ± 1.8 pp, medium confidence
LongCat-Flash-Lite-Sparse45.6%15.8%estimated ± 1.8 pp, medium confidence
Muse Glimmer 30B75.5%18.7%estimated ± 1.8 pp, medium confidence
Muse Spark 1.188.1%19.9%estimated ± 1.8 pp, low confidence
Beam78.7%19.0%estimated ± 1.8 pp, medium confidence
Solar Open 258.2%17.0%estimated ± 1.8 pp, medium confidence
Solar Pro 461.4%17.3%estimated ± 1.8 pp, medium confidence
Step 5 Preview85.6%19.6%estimated ± 1.8 pp, low confidence