Calibration
Claw-Eval → MCP-Tasks
MCP-Tasks is estimated from Claw-Eval with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 1.91595·(x − 0.0469) / (0.0469 + 2.0000 − x), R² = 0.44, cross-validated error 6.4 pp. It is used for 20 estimates.
| Estimated model | Claw-Eval | MCP-Tasks | Source |
|---|---|---|---|
| Claude Opus 4.6 | 70.4% | 93.7% | estimated ± 6.4 pp, low confidence |
| Claude Sonnet 4.6 | 67.8% | 88.3% | estimated ± 6.4 pp, low confidence |
| DeepSeek V3.2 | 40.2% | 41.4% | estimated ± 6.4 pp, low confidence |
| Gemini 3.1 Pro | 57.8% | 69.3% | estimated ± 6.4 pp, low confidence |
| Gemini 3 Flash | 49.2% | 54.8% | estimated ± 6.4 pp, low confidence |
| GLM-5-Turbo | 55.8% | 65.8% | estimated ± 6.4 pp, low confidence |
| GLM-5V-Turbo | 53.8% | 62.4% | estimated ± 6.4 pp, low confidence |
| K-EXAONE 2.0 | 77.7% | 100.0% | estimated ± 6.4 pp, low confidence |
| LLaDA2.2-mini | 57.2% | 68.1% | estimated ± 6.4 pp, low confidence |
| MiMo-V2.5 | 62.3% | 77.5% | estimated ± 6.4 pp, low confidence |
| MiMo-V2-Omni | 45.2% | 48.7% | estimated ± 6.4 pp, low confidence |
| MiMo-V2-Pro | 57.8% | 69.3% | estimated ± 6.4 pp, low confidence |
| MiniMax M2.7 | 48.7% | 54.0% | estimated ± 6.4 pp, low confidence |
| Muse Spark | 63.8% | 80.4% | estimated ± 6.4 pp, low confidence |
| Nemotron 3 Super 100B | 5.5% | 0.8% | estimated ± 6.4 pp, low confidence |
| Ornith-1.0-35B | 69.8% | 92.5% | estimated ± 6.4 pp, low confidence |
| Ornith-1.0-397B | 77.1% | 100.0% | estimated ± 6.4 pp, low confidence |
| Ornith-1.0-9B | 63.1% | 79.0% | estimated ± 6.4 pp, low confidence |
| Qwen3.6-27B | 72.4% | 98.1% | estimated ± 6.4 pp, low confidence |
| Step 3.7 Flash | 67.1% | 86.9% | estimated ± 6.4 pp, low confidence |