benchgap
Calibration

τ²-bench results → AA EnterpriseOps-Gym

AA EnterpriseOps-Gym is estimated from τ²-bench results with a Hill curve fitted on 5 models measured on both: y = 0.0115 + (1.2000 − 0.0115)·x^6.00 / (1.02931^6.00 + x^6.00), R² = 0.93, cross-validated error 3.6 pp. It is used for 126 estimates.

Estimated modelτ²-bench resultsAA EnterpriseOps-GymSource
Agents-A1-4B78.2%20.3%estimated ± 3.6 pp, low confidence
Claude 3 Haiku21.1%1.2%estimated ± 3.6 pp, low confidence
Claude 4.1 Opus Thinking71.4%13.1%estimated ± 3.6 pp, low confidence
Claude 4 Sonnet52.3%3.2%estimated ± 3.6 pp, low confidence
Claude Opus 4.586.3%31.8%estimated ± 3.6 pp, medium confidence
Claude Opus 4.5 Thinking89.5%37.0%estimated ± 3.6 pp, medium confidence
Claude Opus 4.684.8%29.5%estimated ± 3.6 pp, medium confidence
Claude Opus 4.6 (Adaptive)92.1%41.5%estimated ± 3.6 pp, medium confidence
Claude Opus 4.774.0%15.6%estimated ± 3.6 pp, low confidence
Claude Opus 4.7 (Adaptive)88.6%35.5%estimated ± 3.6 pp, medium confidence
Claude Opus 4.894.4%45.5%estimated ± 3.6 pp, medium confidence
Claude Sonnet 4.679.5%22.0%estimated ± 3.6 pp, low confidence
Command A+85.0%29.8%estimated ± 3.6 pp, medium confidence
DeepSeek-R136.5%1.4%estimated ± 3.6 pp, low confidence
DeepSeek V322.8%1.2%estimated ± 3.6 pp, low confidence
DeepSeek V3 032447.1%2.2%estimated ± 3.6 pp, low confidence
DeepSeek V3.134.8%1.3%estimated ± 3.6 pp, low confidence
DeepSeek V3.1 (Reasoning)37.4%1.4%estimated ± 3.6 pp, low confidence
DeepSeek V3.278.9%21.2%estimated ± 3.6 pp, low confidence
Exaone 4.0 1.2B20.5%1.2%estimated ± 3.6 pp, low confidence
Exaone 4.0 32B4.1%1.2%estimated ± 3.6 pp, low confidence
Gemini 2.5 Flash14.9%1.2%estimated ± 3.6 pp, low confidence
Gemini 2.5 Pro54.1%3.6%estimated ± 3.6 pp, low confidence
Gemini 3.1 Pro95.6%47.6%estimated ± 3.6 pp, medium confidence
Gemini 3 Flash43.3%1.8%estimated ± 3.6 pp, low confidence
Gemini 3 Pro87.1%33.1%estimated ± 3.6 pp, medium confidence
Gemma 3 27B10.5%1.2%estimated ± 3.6 pp, low confidence
Gemma 4 12B36.3%1.4%estimated ± 3.6 pp, low confidence
Gemma 4 26B A4B43.6%1.8%estimated ± 3.6 pp, low confidence
Gemma 4 31B59.9%5.6%estimated ± 3.6 pp, low confidence
Gemma 4 E2B20.8%1.2%estimated ± 3.6 pp, low confidence
Gemma 4 E4B20.8%1.2%estimated ± 3.6 pp, low confidence
GLM-4.5-Air46.5%2.2%estimated ± 3.6 pp, low confidence
GLM-4.676.9%18.8%estimated ± 3.6 pp, low confidence
GLM-4.795.9%48.1%estimated ± 3.6 pp, medium confidence
GLM-598.2%52.2%estimated ± 3.6 pp, medium confidence
GLM-5.197.7%51.4%estimated ± 3.6 pp, medium confidence
GLM-5.299.1%53.8%estimated ± 3.6 pp, low confidence
GLM-5-Turbo98.5%52.8%estimated ± 3.6 pp, medium confidence
GLM-5V-Turbo98.5%52.8%estimated ± 3.6 pp, medium confidence
GPT-4.147.1%2.2%estimated ± 3.6 pp, low confidence
GPT-4.1 mini52.9%3.3%estimated ± 3.6 pp, low confidence
GPT-4.1 nano17.3%1.2%estimated ± 3.6 pp, low confidence
GPT-4o25.1%1.2%estimated ± 3.6 pp, low confidence
GPT-5.181.9%25.2%estimated ± 3.6 pp, low confidence
GPT-5.1-Codex83.0%26.8%estimated ± 3.6 pp, low confidence
GPT-5.1-Codex-Max83.0%26.8%estimated ± 3.6 pp, low confidence
GPT-5.284.8%29.5%estimated ± 3.6 pp, medium confidence
GPT-5.2-Codex92.1%41.5%estimated ± 3.6 pp, medium confidence
GPT-5.3 Codex86.0%31.3%estimated ± 3.6 pp, medium confidence
GPT-5.498.9%53.5%estimated ± 3.6 pp, low confidence
GPT-5.4 mini93.4%43.7%estimated ± 3.6 pp, medium confidence
GPT-5.4 nano92.5%42.2%estimated ± 3.6 pp, medium confidence
GPT-5.6 Terra86.3%31.8%estimated ± 3.6 pp, medium confidence
GPT-5 (high)84.8%29.5%estimated ± 3.6 pp, medium confidence
GPT-5 (medium)86.5%32.1%estimated ± 3.6 pp, medium confidence
GPT-OSS 120B65.8%8.7%estimated ± 3.6 pp, low confidence
GPT-OSS 20B60.2%5.7%estimated ± 3.6 pp, low confidence
Granite-4.0-350M13.2%1.2%estimated ± 3.6 pp, low confidence
Granite-4.0-H-1B19.6%1.2%estimated ± 3.6 pp, low confidence
Granite-4.0-H-350M14.6%1.2%estimated ± 3.6 pp, low confidence
Grok 474.9%16.5%estimated ± 3.6 pp, low confidence
Grok 4.1 Fast63.7%7.5%estimated ± 3.6 pp, low confidence
Grok 4.1 Fast (Reasoning)93.3%43.6%estimated ± 3.6 pp, medium confidence
Grok 4.397.7%51.4%estimated ± 3.6 pp, medium confidence
Grok 4 Fast (Reasoning)65.8%8.7%estimated ± 3.6 pp, low confidence
Grok Code Fast 175.7%17.4%estimated ± 3.6 pp, low confidence
K-Exaone74.3%15.9%estimated ± 3.6 pp, low confidence
Kimi K2.695.9%48.1%estimated ± 3.6 pp, medium confidence
Kimi K261.1%6.1%estimated ± 3.6 pp, low confidence
Kimi K2.595.9%48.1%estimated ± 3.6 pp, medium confidence
Kimi K2.5 (Reasoning)95.9%48.1%estimated ± 3.6 pp, medium confidence
Kimi K2.7 Code90.1%38.0%estimated ± 3.6 pp, medium confidence
LFM2.5-8B-A1B88.1%34.6%estimated ± 3.6 pp, medium confidence
LFM2.5-VL-1.6B-Extract8.5%1.2%estimated ± 3.6 pp, low confidence
Ling 2.6 Flash86.0%31.3%estimated ± 3.6 pp, medium confidence
LLaDA2.2-flash80.3%23.1%estimated ± 3.6 pp, low confidence
LLaDA2.2-mini57.5%4.7%estimated ± 3.6 pp, low confidence
Llama 3.1 405B19.0%1.2%estimated ± 3.6 pp, low confidence
Llama 4 Maverick17.8%1.2%estimated ± 3.6 pp, low confidence
Llama 4 Scout15.5%1.2%estimated ± 3.6 pp, low confidence
MiMo-V2.5-Pro94.2%45.1%estimated ± 3.6 pp, medium confidence
MiMo-V2-Flash83.9%28.1%estimated ± 3.6 pp, medium confidence
MiMo-V2-Omni91.2%39.9%estimated ± 3.6 pp, medium confidence
MiMo-V2-Pro95.0%46.6%estimated ± 3.6 pp, medium confidence
MiniMax M2.784.8%29.5%estimated ± 3.6 pp, medium confidence
Mistral Large 230.7%1.2%estimated ± 3.6 pp, low confidence
Mistral Large 324.6%1.2%estimated ± 3.6 pp, low confidence
Mistral Medium 324.3%1.2%estimated ± 3.6 pp, low confidence
Mistral Medium 3.5 128B94.2%45.1%estimated ± 3.6 pp, medium confidence
Mistral Small 441.2%1.6%estimated ± 3.6 pp, low confidence
Mistral Small 4 (Reasoning)41.2%1.6%estimated ± 3.6 pp, low confidence
Muse Spark91.5%40.4%estimated ± 3.6 pp, medium confidence
Nemotron 3 Nano 30B40.9%1.6%estimated ± 3.6 pp, low confidence
Nemotron 3 Nano Omni 30B A3B45.3%2.0%estimated ± 3.6 pp, low confidence
Nemotron 3 Super 100B67.8%10.1%estimated ± 3.6 pp, low confidence
Nemotron Ultra 253B11.4%1.2%estimated ± 3.6 pp, low confidence
North Mini Code37.4%1.4%estimated ± 3.6 pp, low confidence
Nova Pro14.0%1.2%estimated ± 3.6 pp, low confidence
o162.6%6.9%estimated ± 3.6 pp, low confidence
o380.7%23.6%estimated ± 3.6 pp, low confidence
o3-mini28.7%1.2%estimated ± 3.6 pp, low confidence
Phi-40.0%1.2%estimated ± 3.6 pp, low confidence
Qwen3.5-122B-A10B93.6%44.1%estimated ± 3.6 pp, medium confidence
Qwen3.5-27B93.9%44.6%estimated ± 3.6 pp, medium confidence
Qwen3.5-35B-A3B89.2%36.5%estimated ± 3.6 pp, medium confidence
Qwen3.5 397B83.9%28.1%estimated ± 3.6 pp, medium confidence
Qwen3.5 397B (Reasoning)83.9%28.1%estimated ± 3.6 pp, medium confidence
Qwen3.6-27B94.2%45.1%estimated ± 3.6 pp, medium confidence
Qwen3.6-35B-A3B95.3%47.1%estimated ± 3.6 pp, medium confidence
Qwen 3.6 Max (preview)95.9%48.1%estimated ± 3.6 pp, medium confidence
Qwen3.6 Plus97.7%51.4%estimated ± 3.6 pp, medium confidence
Qwen3.7 Max94.7%46.0%estimated ± 3.6 pp, medium confidence
Qwen3.7 Plus93.0%43.0%estimated ± 3.6 pp, medium confidence
Qwen3 Max74.3%15.9%estimated ± 3.6 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct16.4%1.2%estimated ± 3.6 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking21.3%1.2%estimated ± 3.6 pp, low confidence
Sarvam 105B46.8%2.2%estimated ± 3.6 pp, low confidence
Sarvam 30B34.5%1.3%estimated ± 3.6 pp, low confidence
Solar Pro 231.9%1.3%estimated ± 3.6 pp, low confidence
Solar Pro 386.3%31.8%estimated ± 3.6 pp, medium confidence
Step 3.7 Flash98.5%52.8%estimated ± 3.6 pp, medium confidence
Ternary Bonsai 2 27B80.2%22.9%estimated ± 3.6 pp, low confidence
Trinity-Large-Preview90.1%38.0%estimated ± 3.6 pp, medium confidence
Trinity-Large-Thinking90.1%38.0%estimated ± 3.6 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B26.6%1.2%estimated ± 3.6 pp, low confidence