Calibration
MCP Atlas → ResearchClawBench
ResearchClawBench is estimated from MCP Atlas with a linear curve fitted on 12 models measured on both: y = 0.0958·x + 0.1144, R² = 0.49, cross-validated error 1.8 pp. It is used for 19 estimates.
| Estimated model | MCP Atlas | ResearchClawBench | Source |
|---|---|---|---|
| Claude Opus 4.7 (Adaptive) | 77.3% | 18.8% | estimated ± 1.8 pp, medium confidence |
| Claude Opus 5 | 85.8% | 19.7% | estimated ± 1.8 pp, low confidence |
| DeepSeek V4 Flash 0731 | 69.0% | 18.0% | estimated ± 1.8 pp, medium confidence |
| DeepSeek V4 Pro 0813 | 73.6% | 18.5% | estimated ± 1.8 pp, medium confidence |
| GPT-5.4 mini | 57.7% | 17.0% | estimated ± 1.8 pp, medium confidence |
| GPT-5.4 nano | 56.1% | 16.8% | estimated ± 1.8 pp, medium confidence |
| Hy4 preview | 83.7% | 19.5% | estimated ± 1.8 pp, low confidence |
| Inkling | 74.1% | 18.5% | estimated ± 1.8 pp, medium confidence |
| Inkling-Small | 79.6% | 19.1% | estimated ± 1.8 pp, medium confidence |
| Kimi K2.7 Code | 76.0% | 18.7% | estimated ± 1.8 pp, medium confidence |
| Kimi K3 | 84.2% | 19.5% | estimated ± 1.8 pp, low confidence |
| Ling 3.0 Flash | 65.5% | 17.7% | estimated ± 1.8 pp, medium confidence |
| LongCat-Flash-Lite-Sparse | 45.6% | 15.8% | estimated ± 1.8 pp, medium confidence |
| Muse Glimmer 30B | 75.5% | 18.7% | estimated ± 1.8 pp, medium confidence |
| Muse Spark 1.1 | 88.1% | 19.9% | estimated ± 1.8 pp, low confidence |
| Beam | 78.7% | 19.0% | estimated ± 1.8 pp, medium confidence |
| Solar Open 2 | 58.2% | 17.0% | estimated ± 1.8 pp, medium confidence |
| Solar Pro 4 | 61.4% | 17.3% | estimated ± 1.8 pp, medium confidence |
| Step 5 Preview | 85.6% | 19.6% | estimated ± 1.8 pp, low confidence |