benchgap
Calibration

τ²-bench results → HLE w/ tools

HLE w/ tools is estimated from τ²-bench results with a inverse Michaelis–Menten curve fitted on 7 models measured on both: y = 0.69404·(x − 0.1111) / (0.1111 + 2.0000 − x), R² = 0.52, cross-validated error 5.5 pp. It is used for 99 estimates.

Estimated modelτ²-bench resultsHLE w/ toolsSource
Claude Fable 598.5%53.9%estimated ± 5.5 pp, low confidence
GLM-5-Turbo98.5%53.9%estimated ± 5.5 pp, low confidence
GLM-5V-Turbo98.5%53.9%estimated ± 5.5 pp, low confidence
GLM-4.795.9%51.1%estimated ± 5.5 pp, low confidence
Qwen 3.6 Max (preview)95.9%51.1%estimated ± 5.5 pp, low confidence
MiMo-V2-Pro95.0%50.1%estimated ± 5.5 pp, low confidence
Claude Opus 4.894.4%49.5%estimated ± 5.5 pp, low confidence
Mistral Medium 3.5 128B94.2%49.3%estimated ± 5.5 pp, low confidence
Qwen3.6-27B94.2%49.3%estimated ± 5.5 pp, low confidence
Qwen3.5-27B93.9%49.0%estimated ± 5.5 pp, low confidence
Qwen3.5-122B-A10B93.6%48.7%estimated ± 5.5 pp, low confidence
Grok 4.1 Fast (Reasoning)93.3%48.4%estimated ± 5.5 pp, low confidence
GPT-5.2-Codex92.1%47.2%estimated ± 5.5 pp, low confidence
MiMo-V2-Omni91.2%46.4%estimated ± 5.5 pp, low confidence
Kimi K2.7 Code90.1%45.3%estimated ± 5.5 pp, low confidence
Trinity-Large-Preview90.1%45.3%estimated ± 5.5 pp, low confidence
Trinity-Large-Thinking90.1%45.3%estimated ± 5.5 pp, low confidence
Claude Opus 4.5 Thinking89.5%44.7%estimated ± 5.5 pp, low confidence
Qwen3.5-35B-A3B89.2%44.5%estimated ± 5.5 pp, low confidence
MiniMax M388.9%44.2%estimated ± 5.5 pp, low confidence
LFM2.5-8B-A1B88.1%43.4%estimated ± 5.5 pp, low confidence
Gemini 3 Pro87.1%42.5%estimated ± 5.5 pp, low confidence
GPT-5 (medium)86.5%42.0%estimated ± 5.5 pp, low confidence
Solar Pro 386.3%41.8%estimated ± 5.5 pp, low confidence
GPT-5.3 Codex86.0%41.5%estimated ± 5.5 pp, low confidence
Ling 2.6 Flash86.0%41.5%estimated ± 5.5 pp, low confidence
Command A+85.0%40.7%estimated ± 5.5 pp, low confidence
GPT-5.284.8%40.5%estimated ± 5.5 pp, low confidence
GPT-5 (high)84.8%40.5%estimated ± 5.5 pp, low confidence
MiMo-V2-Flash83.9%39.7%estimated ± 5.5 pp, low confidence
Qwen3.5 397B (Reasoning)83.9%39.7%estimated ± 5.5 pp, low confidence
GPT-5.1-Codex83.0%38.9%estimated ± 5.5 pp, low confidence
GPT-5.1-Codex-Max83.0%38.9%estimated ± 5.5 pp, low confidence
GPT-5.181.9%38.0%estimated ± 5.5 pp, low confidence
o380.7%37.0%estimated ± 5.5 pp, low confidence
LLaDA2.2-flash80.3%36.7%estimated ± 5.5 pp, low confidence
Ternary Bonsai 2 27B80.2%36.6%estimated ± 5.5 pp, low confidence
DeepSeek V3.278.9%35.6%estimated ± 5.5 pp, low confidence
Agents-A1-4B78.2%35.0%estimated ± 5.5 pp, low confidence
GLM-4.676.9%34.0%estimated ± 5.5 pp, low confidence
Grok Code Fast 175.7%33.1%estimated ± 5.5 pp, low confidence
Grok 474.9%32.5%estimated ± 5.5 pp, low confidence
K-Exaone74.3%32.1%estimated ± 5.5 pp, low confidence
Qwen3 Max74.3%32.1%estimated ± 5.5 pp, low confidence
Claude Opus 4.774.0%31.8%estimated ± 5.5 pp, low confidence
Claude 4.1 Opus Thinking71.4%30.0%estimated ± 5.5 pp, low confidence
Grok 4 Fast (Reasoning)65.8%26.1%estimated ± 5.5 pp, low confidence
Grok 4.1 Fast63.7%24.8%estimated ± 5.5 pp, low confidence
o162.6%24.1%estimated ± 5.5 pp, low confidence
Kimi K261.1%23.1%estimated ± 5.5 pp, low confidence
Gemma 4 31B59.9%22.4%estimated ± 5.5 pp, low confidence
LLaDA2.2-mini57.5%21.0%estimated ± 5.5 pp, low confidence
Gemini 2.5 Pro54.1%19.0%estimated ± 5.5 pp, low confidence
GPT-4.1 mini52.9%18.3%estimated ± 5.5 pp, low confidence
Claude 4 Sonnet52.3%18.0%estimated ± 5.5 pp, low confidence
DeepSeek V3 032447.1%15.2%estimated ± 5.5 pp, low confidence
GPT-4.147.1%15.2%estimated ± 5.5 pp, low confidence
Sarvam 105B46.8%15.1%estimated ± 5.5 pp, low confidence
GLM-4.5-Air46.5%14.9%estimated ± 5.5 pp, low confidence
Nemotron 3 Nano Omni 30B A3B45.3%14.3%estimated ± 5.5 pp, low confidence
Gemma 4 26B A4B43.6%13.5%estimated ± 5.5 pp, low confidence
Gemini 3 Flash43.3%13.3%estimated ± 5.5 pp, low confidence
Mistral Small 441.2%12.3%estimated ± 5.5 pp, low confidence
Mistral Small 4 (Reasoning)41.2%12.3%estimated ± 5.5 pp, low confidence
Nemotron 3 Nano 30B40.9%12.1%estimated ± 5.5 pp, low confidence
DeepSeek V3.1 (Reasoning)37.4%10.5%estimated ± 5.5 pp, low confidence
North Mini Code37.4%10.5%estimated ± 5.5 pp, low confidence
DeepSeek-R136.5%10.1%estimated ± 5.5 pp, low confidence
Gemma 4 12B36.3%10.0%estimated ± 5.5 pp, low confidence
DeepSeek V3.134.8%9.3%estimated ± 5.5 pp, low confidence
Sarvam 30B34.5%9.2%estimated ± 5.5 pp, low confidence
Solar Pro 231.9%8.1%estimated ± 5.5 pp, low confidence
Mistral Large 230.7%7.5%estimated ± 5.5 pp, low confidence
o3-mini28.7%6.7%estimated ± 5.5 pp, low confidence
Ultravox v0.6 Llama 3.3 70B26.6%5.8%estimated ± 5.5 pp, low confidence
GPT-4o25.1%5.2%estimated ± 5.5 pp, low confidence
Mistral Large 324.6%5.0%estimated ± 5.5 pp, low confidence
Mistral Medium 324.3%4.9%estimated ± 5.5 pp, low confidence
DeepSeek V322.8%4.3%estimated ± 5.5 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking21.3%3.7%estimated ± 5.5 pp, low confidence
Claude 3 Haiku21.1%3.7%estimated ± 5.5 pp, low confidence
Gemma 4 E2B20.8%3.5%estimated ± 5.5 pp, low confidence
Gemma 4 E4B20.8%3.5%estimated ± 5.5 pp, low confidence
Exaone 4.0 1.2B20.5%3.4%estimated ± 5.5 pp, low confidence
Granite-4.0-H-1B19.6%3.1%estimated ± 5.5 pp, low confidence
Llama 3.1 405B19.0%2.9%estimated ± 5.5 pp, low confidence
Llama 4 Maverick17.8%2.4%estimated ± 5.5 pp, low confidence
GPT-4.1 nano17.3%2.2%estimated ± 5.5 pp, low confidence
Qwen3-Omni-30B-A3B-Instruct16.4%1.9%estimated ± 5.5 pp, low confidence
Llama 4 Scout15.5%1.6%estimated ± 5.5 pp, low confidence
Gemini 2.5 Flash14.9%1.3%estimated ± 5.5 pp, low confidence
Granite-4.0-H-350M14.6%1.2%estimated ± 5.5 pp, low confidence
Nova Pro14.0%1.0%estimated ± 5.5 pp, low confidence
Granite-4.0-350M13.2%0.7%estimated ± 5.5 pp, low confidence
Nemotron Ultra 253B11.4%0.1%estimated ± 5.5 pp, low confidence
Exaone 4.0 32B4.1%0.0%estimated ± 5.5 pp, low confidence
Gemma 3 27B10.5%0.0%estimated ± 5.5 pp, low confidence
LFM2.5-VL-1.6B-Extract8.5%0.0%estimated ± 5.5 pp, low confidence
Phi-40.0%0.0%estimated ± 5.5 pp, low confidence