benchgap
Calibration

τ²-bench results → AA ITBench

AA ITBench is estimated from τ²-bench results with a Hill curve fitted on 5 models measured on both: y = 1.0000 + (0.3440 − 1.0000)·x^6.00 / (0.72970^6.00 + x^6.00), R² = 0.77, cross-validated error 3.6 pp. It is used for 107 estimates.

Estimated modelτ²-bench resultsAA ITBenchSource
Agents-A1-4B78.2%60.5%estimated ± 3.6 pp, low confidence
Claude 3 Haiku21.1%100.0%estimated ± 3.6 pp, low confidence
Claude 4.1 Opus Thinking71.4%69.3%estimated ± 3.6 pp, low confidence
Claude 4 Sonnet52.3%92.2%estimated ± 3.6 pp, low confidence
Claude Fable 598.5%43.7%estimated ± 3.6 pp, medium confidence
Claude Opus 4.5 Thinking89.5%49.3%estimated ± 3.6 pp, medium confidence
Claude Opus 4.684.8%53.3%estimated ± 3.6 pp, low confidence
Claude Opus 4.6 (Adaptive)92.1%47.4%estimated ± 3.6 pp, medium confidence
Claude Opus 4.774.0%65.8%estimated ± 3.6 pp, low confidence
Claude Sonnet 4.679.5%58.9%estimated ± 3.6 pp, low confidence
Command A+85.0%53.1%estimated ± 3.6 pp, low confidence
DeepSeek-R136.5%99.0%estimated ± 3.6 pp, low confidence
DeepSeek V322.8%99.9%estimated ± 3.6 pp, low confidence
DeepSeek V3 032447.1%95.6%estimated ± 3.6 pp, low confidence
DeepSeek V3.134.8%99.2%estimated ± 3.6 pp, low confidence
DeepSeek V3.1 (Reasoning)37.4%98.8%estimated ± 3.6 pp, low confidence
DeepSeek V3.278.9%59.6%estimated ± 3.6 pp, low confidence
Exaone 4.0 1.2B20.5%100.0%estimated ± 3.6 pp, low confidence
Exaone 4.0 32B4.1%100.0%estimated ± 3.6 pp, low confidence
Gemini 2.5 Flash14.9%100.0%estimated ± 3.6 pp, low confidence
Gemini 2.5 Pro54.1%90.7%estimated ± 3.6 pp, low confidence
Gemini 3.1 Pro95.6%45.2%estimated ± 3.6 pp, medium confidence
Gemini 3 Flash43.3%97.3%estimated ± 3.6 pp, low confidence
Gemini 3 Pro87.1%51.3%estimated ± 3.6 pp, medium confidence
Gemma 3 27B10.5%100.0%estimated ± 3.6 pp, low confidence
Gemma 4 12B36.3%99.0%estimated ± 3.6 pp, low confidence
Gemma 4 26B A4B43.6%97.1%estimated ± 3.6 pp, low confidence
Gemma 4 31B59.9%84.6%estimated ± 3.6 pp, low confidence
Gemma 4 E2B20.8%100.0%estimated ± 3.6 pp, low confidence
Gemma 4 E4B20.8%100.0%estimated ± 3.6 pp, low confidence
GLM-4.5-Air46.5%95.9%estimated ± 3.6 pp, low confidence
GLM-4.676.9%62.1%estimated ± 3.6 pp, low confidence
GLM-4.795.9%45.1%estimated ± 3.6 pp, medium confidence
GLM-5-Turbo98.5%43.7%estimated ± 3.6 pp, medium confidence
GLM-5V-Turbo98.5%43.7%estimated ± 3.6 pp, medium confidence
GPT-4.147.1%95.6%estimated ± 3.6 pp, low confidence
GPT-4.1 mini52.9%91.7%estimated ± 3.6 pp, low confidence
GPT-4.1 nano17.3%100.0%estimated ± 3.6 pp, low confidence
GPT-4o25.1%99.9%estimated ± 3.6 pp, low confidence
GPT-5.181.9%56.3%estimated ± 3.6 pp, low confidence
GPT-5.1-Codex83.0%55.1%estimated ± 3.6 pp, low confidence
GPT-5.1-Codex-Max83.0%55.1%estimated ± 3.6 pp, low confidence
GPT-5.284.8%53.3%estimated ± 3.6 pp, low confidence
GPT-5.2-Codex92.1%47.4%estimated ± 3.6 pp, medium confidence
GPT-5.3 Codex86.0%52.2%estimated ± 3.6 pp, medium confidence
GPT-5 (high)84.8%53.3%estimated ± 3.6 pp, low confidence
GPT-5 (medium)86.5%51.8%estimated ± 3.6 pp, medium confidence
GPT-OSS 120B65.8%77.1%estimated ± 3.6 pp, low confidence
GPT-OSS 20B60.2%84.3%estimated ± 3.6 pp, low confidence
Granite-4.0-350M13.2%100.0%estimated ± 3.6 pp, low confidence
Granite-4.0-H-1B19.6%100.0%estimated ± 3.6 pp, low confidence
Granite-4.0-H-350M14.6%100.0%estimated ± 3.6 pp, low confidence
Grok 474.9%64.6%estimated ± 3.6 pp, low confidence
Grok 4.1 Fast63.7%79.9%estimated ± 3.6 pp, low confidence
Grok 4.1 Fast (Reasoning)93.3%46.6%estimated ± 3.6 pp, medium confidence
Grok 4.397.7%44.1%estimated ± 3.6 pp, medium confidence
Grok 4 Fast (Reasoning)65.8%77.1%estimated ± 3.6 pp, low confidence
Grok Code Fast 175.7%63.6%estimated ± 3.6 pp, low confidence
K-Exaone74.3%65.4%estimated ± 3.6 pp, low confidence
Kimi K261.1%83.2%estimated ± 3.6 pp, low confidence
Kimi K2.5 (Reasoning)95.9%45.1%estimated ± 3.6 pp, medium confidence
LFM2.5-8B-A1B88.1%50.4%estimated ± 3.6 pp, medium confidence
LFM2.5-VL-1.6B-Extract8.5%100.0%estimated ± 3.6 pp, low confidence
Ling 2.6 Flash86.0%52.2%estimated ± 3.6 pp, medium confidence
LLaDA2.2-mini57.5%87.3%estimated ± 3.6 pp, low confidence
Llama 3.1 405B19.0%100.0%estimated ± 3.6 pp, low confidence
Llama 4 Maverick17.8%100.0%estimated ± 3.6 pp, low confidence
Llama 4 Scout15.5%100.0%estimated ± 3.6 pp, low confidence
MiMo-V2.5-Pro94.2%46.1%estimated ± 3.6 pp, medium confidence
MiMo-V2-Flash83.9%54.2%estimated ± 3.6 pp, low confidence
MiMo-V2-Omni91.2%48.0%estimated ± 3.6 pp, medium confidence
MiMo-V2-Pro95.0%45.6%estimated ± 3.6 pp, medium confidence
MiniMax M2.784.8%53.3%estimated ± 3.6 pp, low confidence
Mistral Large 230.7%99.6%estimated ± 3.6 pp, low confidence
Mistral Large 324.6%99.9%estimated ± 3.6 pp, low confidence
Mistral Medium 324.3%99.9%estimated ± 3.6 pp, low confidence
Mistral Medium 3.5 128B94.2%46.1%estimated ± 3.6 pp, medium confidence
Mistral Small 441.2%97.9%estimated ± 3.6 pp, low confidence
Mistral Small 4 (Reasoning)41.2%97.9%estimated ± 3.6 pp, low confidence
Muse Spark91.5%47.8%estimated ± 3.6 pp, medium confidence
Nemotron 3 Nano 30B40.9%98.0%estimated ± 3.6 pp, low confidence
Nemotron 3 Nano Omni 30B A3B45.3%96.4%estimated ± 3.6 pp, low confidence
Nemotron 3 Super 100B67.8%74.3%estimated ± 3.6 pp, low confidence
Nemotron Ultra 253B11.4%100.0%estimated ± 3.6 pp, low confidence
North Mini Code37.4%98.8%estimated ± 3.6 pp, low confidence
Nova Pro14.0%100.0%estimated ± 3.6 pp, low confidence
o162.6%81.3%estimated ± 3.6 pp, low confidence
o380.7%57.6%estimated ± 3.6 pp, low confidence
o3-mini28.7%99.8%estimated ± 3.6 pp, low confidence
Phi-40.0%100.0%estimated ± 3.6 pp, low confidence
Qwen3.5-122B-A10B93.6%46.4%estimated ± 3.6 pp, medium confidence
Qwen3.5-27B93.9%46.2%estimated ± 3.6 pp, medium confidence
Qwen3.5-35B-A3B89.2%49.5%estimated ± 3.6 pp, medium confidence
Qwen3.5 397B (Reasoning)83.9%54.2%estimated ± 3.6 pp, low confidence
Qwen3.6-27B94.2%46.1%estimated ± 3.6 pp, medium confidence
Qwen 3.6 Max (preview)95.9%45.1%estimated ± 3.6 pp, medium confidence
Qwen3 Max74.3%65.4%estimated ± 3.6 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct16.4%100.0%estimated ± 3.6 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking21.3%100.0%estimated ± 3.6 pp, low confidence
Sarvam 105B46.8%95.7%estimated ± 3.6 pp, low confidence
Sarvam 30B34.5%99.3%estimated ± 3.6 pp, low confidence
Solar Pro 231.9%99.5%estimated ± 3.6 pp, low confidence
Solar Pro 386.3%52.0%estimated ± 3.6 pp, medium confidence
Ternary Bonsai 2 27B80.2%58.1%estimated ± 3.6 pp, low confidence
Trinity-Large-Preview90.1%48.8%estimated ± 3.6 pp, medium confidence
Trinity-Large-Thinking90.1%48.8%estimated ± 3.6 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B26.6%99.8%estimated ± 3.6 pp, low confidence