benchgap
Calibration

HLE w/ tools → WideResearch

WideResearch is estimated from HLE w/ tools with a Michaelis–Menten curve fitted on 6 models measured on both: y = 1.3403·x / (0.35505 + x), R² = 0.95, cross-validated error 3.4 pp. It is used for 16 estimates.

Estimated modelHLE w/ toolsWideResearchSource
Agents-A147.6%76.8%estimated ± 3.4 pp, medium confidence
Apodex 1.156.1%82.1%estimated ± 3.4 pp, medium confidence
Claude Haiku 5.557.4%82.8%estimated ± 3.4 pp, low confidence
Claude Opus 564.7%86.5%estimated ± 3.4 pp, low confidence
Claude Opus 5.567.7%87.9%estimated ± 3.4 pp, low confidence
Claude Sonnet 557.4%82.8%estimated ± 3.4 pp, low confidence
Claude Sonnet 5.564.5%86.4%estimated ± 3.4 pp, low confidence
DeepSeek V4.1 Flash63.9%86.2%estimated ± 3.4 pp, low confidence
DeepSeek V4 Flash 073145.1%75.0%estimated ± 3.4 pp, medium confidence
DeepSeek V4 Pro 081360.0%84.2%estimated ± 3.4 pp, low confidence
GLM-5.362.5%85.5%estimated ± 3.4 pp, low confidence
GLM-5.3-Flash55.3%81.6%estimated ± 3.4 pp, medium confidence
GPT-6 Astra57.2%82.7%estimated ± 3.4 pp, low confidence
Nemotron 3 Ultra37.4%68.8%estimated ± 3.4 pp, medium confidence
Qwen3.7 Max53.5%80.6%estimated ± 3.4 pp, medium confidence
Step 3.7 Flash47.2%76.5%estimated ± 3.4 pp, medium confidence