benchgap
Calibration

τ²-bench results → MCP-Tasks

MCP-Tasks is estimated from τ²-bench results with a linear curve fitted on 5 models measured on both: y = -0.6349·x + 1.2666, R² = 0.33, cross-validated error 8.2 pp. It is used for 91 estimates.

Estimated modelτ²-bench resultsMCP-TasksSource
Agents-A1-4B78.2%77.0%estimated ± 8.2 pp, low confidence
Claude 3 Haiku21.1%100.0%estimated ± 8.2 pp, low confidence
Claude 4.1 Opus Thinking71.4%81.3%estimated ± 8.2 pp, low confidence
Claude 4 Sonnet52.3%93.5%estimated ± 8.2 pp, low confidence
Claude Fable 598.5%64.1%estimated ± 8.2 pp, low confidence
Claude Opus 4.5 Thinking89.5%69.8%estimated ± 8.2 pp, low confidence
Claude Opus 4.6 (Adaptive)92.1%68.2%estimated ± 8.2 pp, low confidence
Claude Opus 4.774.0%79.7%estimated ± 8.2 pp, low confidence
Command A+85.0%72.7%estimated ± 8.2 pp, low confidence
DeepSeek-R136.5%100.0%estimated ± 8.2 pp, low confidence
DeepSeek V322.8%100.0%estimated ± 8.2 pp, low confidence
DeepSeek V3 032447.1%96.8%estimated ± 8.2 pp, low confidence
DeepSeek V3.134.8%100.0%estimated ± 8.2 pp, low confidence
DeepSeek V3.1 (Reasoning)37.4%100.0%estimated ± 8.2 pp, low confidence
Exaone 4.0 1.2B20.5%100.0%estimated ± 8.2 pp, low confidence
Exaone 4.0 32B4.1%100.0%estimated ± 8.2 pp, low confidence
Gemini 2.5 Flash14.9%100.0%estimated ± 8.2 pp, low confidence
Gemini 2.5 Pro54.1%92.3%estimated ± 8.2 pp, low confidence
Gemini 3 Pro87.1%71.4%estimated ± 8.2 pp, low confidence
Gemma 3 27B10.5%100.0%estimated ± 8.2 pp, low confidence
Gemma 4 12B36.3%100.0%estimated ± 8.2 pp, low confidence
Gemma 4 26B A4B43.6%99.0%estimated ± 8.2 pp, low confidence
Gemma 4 31B59.9%88.6%estimated ± 8.2 pp, low confidence
Gemma 4 E2B20.8%100.0%estimated ± 8.2 pp, low confidence
Gemma 4 E4B20.8%100.0%estimated ± 8.2 pp, low confidence
GLM-4.5-Air46.5%97.1%estimated ± 8.2 pp, low confidence
GLM-4.676.9%77.8%estimated ± 8.2 pp, low confidence
GLM-4.795.9%65.8%estimated ± 8.2 pp, low confidence
GPT-4.147.1%96.8%estimated ± 8.2 pp, low confidence
GPT-4.1 mini52.9%93.1%estimated ± 8.2 pp, low confidence
GPT-4.1 nano17.3%100.0%estimated ± 8.2 pp, low confidence
GPT-4o25.1%100.0%estimated ± 8.2 pp, low confidence
GPT-5.181.9%74.7%estimated ± 8.2 pp, low confidence
GPT-5.1-Codex83.0%74.0%estimated ± 8.2 pp, low confidence
GPT-5.1-Codex-Max83.0%74.0%estimated ± 8.2 pp, low confidence
GPT-5.284.8%72.8%estimated ± 8.2 pp, low confidence
GPT-5.2-Codex92.1%68.2%estimated ± 8.2 pp, low confidence
GPT-5.3 Codex86.0%72.1%estimated ± 8.2 pp, low confidence
GPT-5 (high)84.8%72.8%estimated ± 8.2 pp, low confidence
GPT-5 (medium)86.5%71.7%estimated ± 8.2 pp, low confidence
GPT-OSS 120B65.8%84.9%estimated ± 8.2 pp, low confidence
GPT-OSS 20B60.2%88.4%estimated ± 8.2 pp, low confidence
Granite-4.0-350M13.2%100.0%estimated ± 8.2 pp, low confidence
Granite-4.0-H-1B19.6%100.0%estimated ± 8.2 pp, low confidence
Granite-4.0-H-350M14.6%100.0%estimated ± 8.2 pp, low confidence
Grok 474.9%79.1%estimated ± 8.2 pp, low confidence
Grok 4.1 Fast63.7%86.2%estimated ± 8.2 pp, low confidence
Grok 4.1 Fast (Reasoning)93.3%67.4%estimated ± 8.2 pp, low confidence
Grok 4.397.7%64.6%estimated ± 8.2 pp, low confidence
Grok 4 Fast (Reasoning)65.8%84.9%estimated ± 8.2 pp, low confidence
Grok Code Fast 175.7%78.6%estimated ± 8.2 pp, low confidence
K-Exaone74.3%79.5%estimated ± 8.2 pp, low confidence
Kimi K261.1%87.9%estimated ± 8.2 pp, low confidence
Kimi K2.5 (Reasoning)95.9%65.8%estimated ± 8.2 pp, low confidence
LFM2.5-8B-A1B88.1%70.7%estimated ± 8.2 pp, low confidence
LFM2.5-VL-1.6B-Extract8.5%100.0%estimated ± 8.2 pp, low confidence
Ling 2.6 Flash86.0%72.1%estimated ± 8.2 pp, low confidence
Llama 3.1 405B19.0%100.0%estimated ± 8.2 pp, low confidence
Llama 4 Maverick17.8%100.0%estimated ± 8.2 pp, low confidence
Llama 4 Scout15.5%100.0%estimated ± 8.2 pp, low confidence
MiMo-V2-Flash83.9%73.4%estimated ± 8.2 pp, low confidence
Mistral Large 230.7%100.0%estimated ± 8.2 pp, low confidence
Mistral Large 324.6%100.0%estimated ± 8.2 pp, low confidence
Mistral Medium 324.3%100.0%estimated ± 8.2 pp, low confidence
Mistral Small 441.2%100.0%estimated ± 8.2 pp, low confidence
Mistral Small 4 (Reasoning)41.2%100.0%estimated ± 8.2 pp, low confidence
Nemotron 3 Nano 30B40.9%100.0%estimated ± 8.2 pp, low confidence
Nemotron 3 Nano Omni 30B A3B45.3%97.9%estimated ± 8.2 pp, low confidence
Nemotron Ultra 253B11.4%100.0%estimated ± 8.2 pp, low confidence
North Mini Code37.4%100.0%estimated ± 8.2 pp, low confidence
Nova Pro14.0%100.0%estimated ± 8.2 pp, low confidence
o162.6%86.9%estimated ± 8.2 pp, low confidence
o380.7%75.4%estimated ± 8.2 pp, low confidence
o3-mini28.7%100.0%estimated ± 8.2 pp, low confidence
Phi-40.0%100.0%estimated ± 8.2 pp, low confidence
Qwen3.5-122B-A10B93.6%67.2%estimated ± 8.2 pp, low confidence
Qwen3.5-27B93.9%67.0%estimated ± 8.2 pp, low confidence
Qwen3.5-35B-A3B89.2%70.0%estimated ± 8.2 pp, low confidence
Qwen3.5 397B (Reasoning)83.9%73.4%estimated ± 8.2 pp, low confidence
Qwen 3.6 Max (preview)95.9%65.8%estimated ± 8.2 pp, low confidence
Qwen3 Max74.3%79.5%estimated ± 8.2 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct16.4%100.0%estimated ± 8.2 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking21.3%100.0%estimated ± 8.2 pp, low confidence
Sarvam 105B46.8%96.9%estimated ± 8.2 pp, low confidence
Sarvam 30B34.5%100.0%estimated ± 8.2 pp, low confidence
Solar Pro 231.9%100.0%estimated ± 8.2 pp, low confidence
Solar Pro 386.3%71.9%estimated ± 8.2 pp, low confidence
Ternary Bonsai 2 27B80.2%75.7%estimated ± 8.2 pp, low confidence
Trinity-Large-Preview90.1%69.5%estimated ± 8.2 pp, low confidence
Trinity-Large-Thinking90.1%69.5%estimated ± 8.2 pp, low confidence
Ultravox v0.6 Llama 3.3 70B26.6%100.0%estimated ± 8.2 pp, low confidence