benchgap
Calibration

τ²-bench results → APEX-Agents-AA

APEX-Agents-AA is estimated from τ²-bench results with a offset logistic curve fitted on 22 models measured on both: y = 0.0165 + (0.2527 − 0.0165) / (1 + exp(−200.00·(x − 0.8481))), R² = 0.39, cross-validated error 11.4 pp. It is used for 48 estimates.

Estimated modelτ²-bench resultsAPEX-Agents-AASource
Claude 3 Haiku21.1%1.6%estimated ± 11.4 pp, low confidence
Claude 4.1 Opus Thinking71.4%1.6%estimated ± 11.4 pp, low confidence
Claude Opus 4.5 Thinking89.5%25.3%estimated ± 11.4 pp, low confidence
DeepSeek-R136.5%1.6%estimated ± 11.4 pp, low confidence
DeepSeek V3.134.8%1.6%estimated ± 11.4 pp, low confidence
DeepSeek V3.1 (Reasoning)37.4%1.6%estimated ± 11.4 pp, low confidence
DeepSeek V3.278.9%1.6%estimated ± 11.4 pp, low confidence
Exaone 4.0 1.2B20.5%1.6%estimated ± 11.4 pp, low confidence
Exaone 4.0 32B4.1%1.6%estimated ± 11.4 pp, low confidence
Gemini 2.5 Flash14.9%1.6%estimated ± 11.4 pp, low confidence
GLM-4.5-Air46.5%1.6%estimated ± 11.4 pp, low confidence
GLM-4.676.9%1.6%estimated ± 11.4 pp, low confidence
GLM-5-Turbo98.5%25.3%estimated ± 11.4 pp, low confidence
GLM-5V-Turbo98.5%25.3%estimated ± 11.4 pp, low confidence
GPT-4.147.1%1.6%estimated ± 11.4 pp, low confidence
GPT-5.1-Codex-Max83.0%2.3%estimated ± 11.4 pp, low confidence
GPT-5 (medium)86.5%24.5%estimated ± 11.4 pp, low confidence
Granite-4.0-350M13.2%1.6%estimated ± 11.4 pp, low confidence
Granite-4.0-H-1B19.6%1.6%estimated ± 11.4 pp, low confidence
Granite-4.0-H-350M14.6%1.6%estimated ± 11.4 pp, low confidence
Grok 474.9%1.6%estimated ± 11.4 pp, low confidence
Grok 4.1 Fast63.7%1.6%estimated ± 11.4 pp, low confidence
Grok 4.1 Fast (Reasoning)93.3%25.3%estimated ± 11.4 pp, low confidence
Grok 4 Fast (Reasoning)65.8%1.6%estimated ± 11.4 pp, low confidence
Grok Code Fast 175.7%1.6%estimated ± 11.4 pp, low confidence
Kimi K261.1%1.6%estimated ± 11.4 pp, low confidence
LFM2.5-8B-A1B88.1%25.2%estimated ± 11.4 pp, low confidence
LFM2.5-VL-1.6B-Extract8.5%1.6%estimated ± 11.4 pp, low confidence
LLaDA2.2-mini57.5%1.6%estimated ± 11.4 pp, low confidence
Llama 3.1 405B19.0%1.6%estimated ± 11.4 pp, low confidence
MiMo-V2-Omni91.2%25.3%estimated ± 11.4 pp, low confidence
Mistral Large 230.7%1.6%estimated ± 11.4 pp, low confidence
Mistral Medium 324.3%1.6%estimated ± 11.4 pp, low confidence
Nemotron Ultra 253B11.4%1.6%estimated ± 11.4 pp, low confidence
Nova Pro14.0%1.6%estimated ± 11.4 pp, low confidence
o162.6%1.6%estimated ± 11.4 pp, low confidence
o380.7%1.7%estimated ± 11.4 pp, low confidence
o3-mini28.7%1.6%estimated ± 11.4 pp, low confidence
Phi-40.0%1.6%estimated ± 11.4 pp, low confidence
Qwen3.5 397B (Reasoning)83.9%4.9%estimated ± 11.4 pp, low confidence
Qwen 3.6 Max (preview)95.9%25.3%estimated ± 11.4 pp, low confidence
Qwen3 Max74.3%1.6%estimated ± 11.4 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct16.4%1.6%estimated ± 11.4 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking21.3%1.6%estimated ± 11.4 pp, low confidence
Sarvam 105B46.8%1.6%estimated ± 11.4 pp, low confidence
Sarvam 30B34.5%1.6%estimated ± 11.4 pp, low confidence
Solar Pro 231.9%1.6%estimated ± 11.4 pp, low confidence
Ternary Bonsai 2 27B80.2%1.6%estimated ± 11.4 pp, low confidence