benchgap
Calibration

AA-HLE → HLE w/o tools

HLE w/o tools is estimated from AA-HLE with a inverse Michaelis–Menten curve fitted on 28 models measured on both: y = 1.53498·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.86, cross-validated error 5.7 pp. It is used for 94 estimates.

Estimated modelAA-HLEHLE w/o toolsSource
Apodex 1.1 Mini34.1%31.6%estimated ± 5.7 pp, medium confidence
Celeris-16.8%5.4%estimated ± 5.7 pp, low confidence
Claude 3 Haiku4.1%3.2%estimated ± 5.7 pp, low confidence
Claude 3 Opus2.8%2.2%estimated ± 5.7 pp, low confidence
Claude 4.1 Opus Thinking12.5%10.2%estimated ± 5.7 pp, medium confidence
Claude 4 Sonnet4.3%3.4%estimated ± 5.7 pp, low confidence
Claude Opus 4.5 Thinking30.1%27.2%estimated ± 5.7 pp, medium confidence
Claude Opus 4.6 (Adaptive)39.9%38.3%estimated ± 5.7 pp, medium confidence
Command A+12.0%9.8%estimated ± 5.7 pp, medium confidence
DeepSeek-R115.8%13.2%estimated ± 5.7 pp, medium confidence
DeepSeek R1 Distill Qwen 32B4.6%3.6%estimated ± 5.7 pp, low confidence
DeepSeek V3 03244.7%3.7%estimated ± 5.7 pp, low confidence
DeepSeek V3.16.7%5.3%estimated ± 5.7 pp, low confidence
DeepSeek V3.1 (Reasoning)14.3%11.8%estimated ± 5.7 pp, medium confidence
DeepSeek V3.211.2%9.1%estimated ± 5.7 pp, medium confidence
Exaone 4.0 1.2B5.7%4.5%estimated ± 5.7 pp, low confidence
Exaone 4.0 32B5.0%3.9%estimated ± 5.7 pp, low confidence
Gemini 1.0 Pro4.2%3.3%estimated ± 5.7 pp, low confidence
Gemini 1.5 Pro4.6%3.6%estimated ± 5.7 pp, low confidence
Gemini 2.5 Flash4.7%3.7%estimated ± 5.7 pp, low confidence
Gemini 3 Pro39.7%38.0%estimated ± 5.7 pp, medium confidence
Gemini 4 Argon57.1%61.3%estimated ± 5.7 pp, medium confidence
Gemma 3 27B4.4%3.5%estimated ± 5.7 pp, low confidence
GLM-4.5-Air7.0%5.6%estimated ± 5.7 pp, low confidence
GLM-5-Turbo27.8%24.8%estimated ± 5.7 pp, medium confidence
GLM-5V-Turbo17.1%14.4%estimated ± 5.7 pp, medium confidence
GPT-4 Turbo3.1%2.4%estimated ± 5.7 pp, low confidence
GPT-4o2.4%1.9%estimated ± 5.7 pp, low confidence
GPT-4o mini4.2%3.3%estimated ± 5.7 pp, low confidence
GPT-5.128.5%25.5%estimated ± 5.7 pp, medium confidence
GPT-5.1-Codex25.7%22.6%estimated ± 5.7 pp, medium confidence
GPT-5.1-Codex-Max25.7%22.6%estimated ± 5.7 pp, medium confidence
GPT-5.2-Codex35.7%33.4%estimated ± 5.7 pp, medium confidence
GPT-5.3 Codex42.5%41.4%estimated ± 5.7 pp, medium confidence
GPT-5 (high)28.5%25.5%estimated ± 5.7 pp, medium confidence
GPT-5 (medium)25.4%22.3%estimated ± 5.7 pp, medium confidence
GPT-OSS 120B19.6%16.7%estimated ± 5.7 pp, medium confidence
GPT-OSS 20B11.0%8.9%estimated ± 5.7 pp, medium confidence
Granite-4.0-350M5.5%4.3%estimated ± 5.7 pp, low confidence
Granite-4.0-H-1B5.0%3.9%estimated ± 5.7 pp, low confidence
Granite-4.0-H-350M6.4%5.1%estimated ± 5.7 pp, low confidence
Grok 426.7%23.6%estimated ± 5.7 pp, medium confidence
Grok 4.1 Fast5.1%4.0%estimated ± 5.7 pp, low confidence
Grok 4.1 Fast (Reasoning)19.3%16.4%estimated ± 5.7 pp, medium confidence
Grok 4.743.1%42.2%estimated ± 5.7 pp, medium confidence
Grok 4 Fast (Reasoning)19.1%16.2%estimated ± 5.7 pp, medium confidence
Grok Code Fast 18.0%6.4%estimated ± 5.7 pp, low confidence
Hy333.5%30.9%estimated ± 5.7 pp, medium confidence
K-Exaone13.9%11.5%estimated ± 5.7 pp, medium confidence
Kimi K27.4%5.9%estimated ± 5.7 pp, low confidence
Kimi K2.7 Code35.0%32.6%estimated ± 5.7 pp, medium confidence
LFM2.5-2.6B6.2%4.9%estimated ± 5.7 pp, low confidence
LFM2.5-8B-A1B6.9%5.5%estimated ± 5.7 pp, low confidence
LFM2.5-VL-1.6B-Extract5.1%4.0%estimated ± 5.7 pp, low confidence
Ling 3.0 Flash VL22.0%19.0%estimated ± 5.7 pp, medium confidence
Ling 3.0 Tiny9.3%7.5%estimated ± 5.7 pp, low confidence
Ling 3.1 Flash39.4%37.7%estimated ± 5.7 pp, medium confidence
Llama 3.1 405B4.0%3.1%estimated ± 5.7 pp, low confidence
Llama 4 Maverick4.9%3.9%estimated ± 5.7 pp, low confidence
Llama 4 Scout3.8%3.0%estimated ± 5.7 pp, low confidence
MiMo-V2.6-Flash35.1%32.7%estimated ± 5.7 pp, medium confidence
MiMo-V2.6-Pro49.4%50.4%estimated ± 5.7 pp, medium confidence
MiMo-V2-Omni22.1%19.1%estimated ± 5.7 pp, medium confidence
MiMo-V2-Pro30.4%27.5%estimated ± 5.7 pp, medium confidence
Mistral Large 23.3%2.6%estimated ± 5.7 pp, low confidence
Mistral Large 34.2%3.3%estimated ± 5.7 pp, low confidence
Mistral Large 435.0%32.6%estimated ± 5.7 pp, medium confidence
Mistral Medium 34.1%3.2%estimated ± 5.7 pp, low confidence
Mistral Small 49.9%8.0%estimated ± 5.7 pp, low confidence
Mistral Small 4 (Reasoning)9.9%8.0%estimated ± 5.7 pp, low confidence
Muse Glimmer 30B22.0%19.0%estimated ± 5.7 pp, medium confidence
Muse Spark 1.348.7%49.4%estimated ± 5.7 pp, medium confidence
Nemotron 3 Nano 30B11.4%9.3%estimated ± 5.7 pp, medium confidence
Nemotron 3 Super 100B20.8%17.8%estimated ± 5.7 pp, medium confidence
Nemotron Ultra 253B7.4%5.9%estimated ± 5.7 pp, low confidence
North Mini Code11.1%9.0%estimated ± 5.7 pp, medium confidence
Nova Pro3.2%2.5%estimated ± 5.7 pp, low confidence
o320.1%17.2%estimated ± 5.7 pp, medium confidence
Phi-43.8%3.0%estimated ± 5.7 pp, low confidence
Phi-4 Multimodal Instruct5.0%3.9%estimated ± 5.7 pp, low confidence
Quasar 438B18.7%15.8%estimated ± 5.7 pp, medium confidence
Qwen2.5 Coder 32B Instruct3.5%2.7%estimated ± 5.7 pp, low confidence
Qwen3.5 397B (Reasoning)19.8%16.9%estimated ± 5.7 pp, medium confidence
Qwen 3.6 Max (preview)30.8%27.9%estimated ± 5.7 pp, medium confidence
Qwen3.8 Max Preview43.1%42.2%estimated ± 5.7 pp, medium confidence
Qwen3 Max11.9%9.7%estimated ± 5.7 pp, medium confidence
Qwen3-Omni-30B-A3B-Instruct4.6%3.6%estimated ± 5.7 pp, low confidence
Qwen3-Omni-30B-A3B-Thinking7.5%6.0%estimated ± 5.7 pp, low confidence
Sarvam 105B11.0%8.9%estimated ± 5.7 pp, medium confidence
Sarvam 30B7.5%6.0%estimated ± 5.7 pp, low confidence
Solar Pro 23.7%2.9%estimated ± 5.7 pp, low confidence
Solar Pro 310.3%8.3%estimated ± 5.7 pp, low confidence
Step 3.7 Flash21.4%18.4%estimated ± 5.7 pp, medium confidence
Ultravox v0.6 Llama 3.3 70B3.6%2.8%estimated ± 5.7 pp, low confidence