benchgap
Calibration

MCP Atlas → Claw-Eval

Claw-Eval is estimated from MCP Atlas with a linear curve fitted on 16 models measured on both: y = 0.3266·x + 0.4507, R² = 0.53, cross-validated error 5.7 pp. It is used for 13 estimates.

Estimated modelMCP AtlasClaw-EvalSource
Claude Opus 585.8%73.1%estimated ± 5.7 pp, low confidence
DeepSeek V4 Flash 073169.0%67.6%estimated ± 5.7 pp, medium confidence
DeepSeek V4 Pro 081373.6%69.1%estimated ± 5.7 pp, medium confidence
GPT-5.4 mini57.7%63.9%estimated ± 5.7 pp, medium confidence
GPT-5.4 nano56.1%63.4%estimated ± 5.7 pp, medium confidence
Inkling74.1%69.3%estimated ± 5.7 pp, medium confidence
Inkling-Small79.6%71.1%estimated ± 5.7 pp, medium confidence
Kimi K2.7 Code76.0%69.9%estimated ± 5.7 pp, medium confidence
LongCat-Flash-Lite-Sparse45.6%60.0%estimated ± 5.7 pp, medium confidence
Muse Glimmer 30B75.5%69.7%estimated ± 5.7 pp, medium confidence
Beam78.7%70.8%estimated ± 5.7 pp, medium confidence
Solar Open 258.2%64.1%estimated ± 5.7 pp, medium confidence
Solar Pro 461.4%65.1%estimated ± 5.7 pp, medium confidence