benchgap
Calibration

MCP Atlas → JobBench

JobBench is estimated from MCP Atlas with a inverse Michaelis–Menten curve fitted on 9 models measured on both: y = 0.75626·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.88, cross-validated error 6.8 pp. It is used for 14 estimates.

Estimated modelMCP AtlasJobBenchSource
Gemini 3.5 Flash83.6%54.3%estimated ± 6.8 pp, medium confidence
GLM-5.171.8%42.4%estimated ± 6.8 pp, medium confidence
GLM-5.276.8%47.1%estimated ± 6.8 pp, medium confidence
GPT-5.4 mini57.7%30.7%estimated ± 6.8 pp, medium confidence
GPT-5.4 nano56.1%29.5%estimated ± 6.8 pp, medium confidence
Inkling74.1%44.5%estimated ± 6.8 pp, medium confidence
Kimi K2.7 Code76.0%46.4%estimated ± 6.8 pp, medium confidence
LLaDA2.2-flash46.2%22.7%estimated ± 6.8 pp, medium confidence
LongCat-Flash-Lite-Sparse45.6%22.3%estimated ± 6.8 pp, medium confidence
Muse Glimmer 30B75.5%45.9%estimated ± 6.8 pp, medium confidence
Qwen3.7 Max76.4%46.7%estimated ± 6.8 pp, medium confidence
Beam78.7%49.1%estimated ± 6.8 pp, medium confidence
Solar Open 258.2%31.0%estimated ± 6.8 pp, medium confidence
Solar Pro 461.4%33.5%estimated ± 6.8 pp, medium confidence