benchgap
Calibration

Claw-Eval → MCP-Tasks

MCP-Tasks is estimated from Claw-Eval with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 1.91595·(x − 0.0469) / (0.0469 + 2.0000 − x), R² = 0.44, cross-validated error 6.4 pp. It is used for 20 estimates.

Estimated modelClaw-EvalMCP-TasksSource
Claude Opus 4.670.4%93.7%estimated ± 6.4 pp, low confidence
Claude Sonnet 4.667.8%88.3%estimated ± 6.4 pp, low confidence
DeepSeek V3.240.2%41.4%estimated ± 6.4 pp, low confidence
Gemini 3.1 Pro57.8%69.3%estimated ± 6.4 pp, low confidence
Gemini 3 Flash49.2%54.8%estimated ± 6.4 pp, low confidence
GLM-5-Turbo55.8%65.8%estimated ± 6.4 pp, low confidence
GLM-5V-Turbo53.8%62.4%estimated ± 6.4 pp, low confidence
K-EXAONE 2.077.7%100.0%estimated ± 6.4 pp, low confidence
LLaDA2.2-mini57.2%68.1%estimated ± 6.4 pp, low confidence
MiMo-V2.562.3%77.5%estimated ± 6.4 pp, low confidence
MiMo-V2-Omni45.2%48.7%estimated ± 6.4 pp, low confidence
MiMo-V2-Pro57.8%69.3%estimated ± 6.4 pp, low confidence
MiniMax M2.748.7%54.0%estimated ± 6.4 pp, low confidence
Muse Spark63.8%80.4%estimated ± 6.4 pp, low confidence
Nemotron 3 Super 100B5.5%0.8%estimated ± 6.4 pp, low confidence
Ornith-1.0-35B69.8%92.5%estimated ± 6.4 pp, low confidence
Ornith-1.0-397B77.1%100.0%estimated ± 6.4 pp, low confidence
Ornith-1.0-9B63.1%79.0%estimated ± 6.4 pp, low confidence
Qwen3.6-27B72.4%98.1%estimated ± 6.4 pp, low confidence
Step 3.7 Flash67.1%86.9%estimated ± 6.4 pp, low confidence