Calibration
MCP Atlas → APEX-Agents
APEX-Agents is estimated from MCP Atlas with a linear curve fitted on 5 models measured on both: y = 0.8014·x + -0.3024, R² = 1.00, cross-validated error 0.6 pp. It is used for 34 estimates.
| Estimated model | MCP Atlas | APEX-Agents | Source |
|---|---|---|---|
| Claude Opus 4.5 | 42.3% | 3.7% | estimated ± 0.6 pp, low confidence |
| Claude Opus 4.7 (Adaptive) | 77.3% | 31.7% | estimated ± 0.6 pp, medium confidence |
| Claude Opus 4.8 | 82.2% | 35.6% | estimated ± 0.6 pp, medium confidence |
| Claude Opus 5 | 85.8% | 38.5% | estimated ± 0.6 pp, low confidence |
| DeepSeek V4 Flash 0731 | 69.0% | 25.1% | estimated ± 0.6 pp, medium confidence |
| DeepSeek V4 Pro 0813 | 73.6% | 28.7% | estimated ± 0.6 pp, medium confidence |
| Gemini 3.5 Flash | 83.6% | 36.8% | estimated ± 0.6 pp, medium confidence |
| GLM-5 | 31.1% | 0.0% | estimated ± 0.6 pp, low confidence |
| GLM-5.1 | 71.8% | 27.3% | estimated ± 0.6 pp, medium confidence |
| GLM-5.2 | 76.8% | 31.3% | estimated ± 0.6 pp, medium confidence |
| GPT-5.4 | 70.6% | 26.3% | estimated ± 0.6 pp, medium confidence |
| GPT-5.4 mini | 57.7% | 16.0% | estimated ± 0.6 pp, low confidence |
| GPT-5.4 nano | 56.1% | 14.7% | estimated ± 0.6 pp, low confidence |
| GPT-5.5 | 75.3% | 30.1% | estimated ± 0.6 pp, medium confidence |
| Inkling | 74.1% | 29.1% | estimated ± 0.6 pp, medium confidence |
| Inkling-Small | 79.6% | 33.6% | estimated ± 0.6 pp, medium confidence |
| Kimi K2.6 | 55.9% | 14.6% | estimated ± 0.6 pp, low confidence |
| Kimi K2.5 | 29.5% | 0.0% | estimated ± 0.6 pp, low confidence |
| Kimi K2.7 Code | 76.0% | 30.7% | estimated ± 0.6 pp, medium confidence |
| Ling 3.0 Flash | 65.5% | 22.3% | estimated ± 0.6 pp, medium confidence |
| LLaDA2.2-flash | 46.2% | 6.8% | estimated ± 0.6 pp, low confidence |
| LongCat-Flash-Lite-Sparse | 45.6% | 6.3% | estimated ± 0.6 pp, low confidence |
| MiniMax M3 | 74.2% | 29.2% | estimated ± 0.6 pp, medium confidence |
| Muse Glimmer 30B | 75.5% | 30.3% | estimated ± 0.6 pp, medium confidence |
| Muse Spark 1.1 | 88.1% | 40.4% | estimated ± 0.6 pp, low confidence |
| Ornith-1.5-35B-A3B | 70.2% | 26.0% | estimated ± 0.6 pp, medium confidence |
| Ornith-1.5-397B | 80.0% | 33.9% | estimated ± 0.6 pp, medium confidence |
| Ornith-1.5-9B | 54.2% | 13.2% | estimated ± 0.6 pp, low confidence |
| Qwen3.5 397B | 46.1% | 6.7% | estimated ± 0.6 pp, low confidence |
| Qwen3.6-35B-A3B | 62.8% | 20.1% | estimated ± 0.6 pp, medium confidence |
| Qwen3.6 Plus | 48.2% | 8.4% | estimated ± 0.6 pp, low confidence |
| Qwen3.7 Max | 76.4% | 31.0% | estimated ± 0.6 pp, medium confidence |
| Qwen3.7 Plus | 73.2% | 28.4% | estimated ± 0.6 pp, medium confidence |
| Beam | 78.7% | 32.8% | estimated ± 0.6 pp, medium confidence |