benchgap
Calibration

MCP Atlas → VITA-Bench

VITA-Bench is estimated from MCP Atlas with a Michaelis–Menten curve fitted on 7 models measured on both: y = 1.3829·x / (1.49552 + x), R² = 0.40, cross-validated error 9.9 pp. It is used for 30 estimates.

Estimated modelMCP AtlasVITA-BenchSource
Claude Opus 4.7 (Adaptive)77.3%47.1%estimated ± 9.9 pp, low confidence
Claude Opus 4.882.2%49.0%estimated ± 9.9 pp, low confidence
Claude Opus 585.8%50.4%estimated ± 9.9 pp, low confidence
DeepSeek V4 Flash 073169.0%43.7%estimated ± 9.9 pp, low confidence
DeepSeek V4 Pro 081373.6%45.6%estimated ± 9.9 pp, low confidence
Gemini 3.5 Flash83.6%49.6%estimated ± 9.9 pp, low confidence
GLM-5.171.8%44.9%estimated ± 9.9 pp, low confidence
GLM-5.276.8%46.9%estimated ± 9.9 pp, low confidence
GPT-5.470.6%44.3%estimated ± 9.9 pp, low confidence
GPT-5.4 mini57.7%38.5%estimated ± 9.9 pp, low confidence
GPT-5.4 nano56.1%37.7%estimated ± 9.9 pp, low confidence
GPT-5.575.3%46.3%estimated ± 9.9 pp, low confidence
Hy4 preview83.7%49.6%estimated ± 9.9 pp, low confidence
Inkling74.1%45.8%estimated ± 9.9 pp, low confidence
Inkling-Small79.6%48.0%estimated ± 9.9 pp, low confidence
Kimi K2.655.9%37.6%estimated ± 9.9 pp, low confidence
Kimi K2.7 Code76.0%46.6%estimated ± 9.9 pp, low confidence
Kimi K384.2%49.8%estimated ± 9.9 pp, low confidence
Ling 3.0 Flash65.5%42.1%estimated ± 9.9 pp, low confidence
LLaDA2.2-flash46.2%32.6%estimated ± 9.9 pp, low confidence
MiniMax M374.2%45.9%estimated ± 9.9 pp, low confidence
Muse Glimmer 30B75.5%46.4%estimated ± 9.9 pp, low confidence
Muse Spark 1.188.1%51.3%estimated ± 9.9 pp, low confidence
Ornith-1.5-35B-A3B70.2%44.2%estimated ± 9.9 pp, low confidence
Ornith-1.5-397B80.0%48.2%estimated ± 9.9 pp, low confidence
Ornith-1.5-9B54.2%36.8%estimated ± 9.9 pp, low confidence
Beam78.7%47.7%estimated ± 9.9 pp, low confidence
Solar Open 258.2%38.7%estimated ± 9.9 pp, low confidence
Solar Pro 461.4%40.3%estimated ± 9.9 pp, low confidence
Step 5 Preview85.6%50.3%estimated ± 9.9 pp, low confidence