benchgap
Calibration

JobBench → WideResearch

WideResearch is estimated from JobBench with a linear curve fitted on 5 models measured on both: y = 0.2183·x + 0.7035, R² = 0.98, cross-validated error 1.2 pp. It is used for 25 estimates.

Estimated modelJobBenchWideResearchSource
Claude 4.1 Opus21.9%75.1%estimated ± 1.2 pp, medium confidence
Claude 4 Sonnet18.4%74.4%estimated ± 1.2 pp, medium confidence
Claude Haiku 4.516.0%73.8%estimated ± 1.2 pp, medium confidence
Claude Opus 4.636.7%78.4%estimated ± 1.2 pp, medium confidence
Claude Opus 4.7 (Adaptive)45.9%80.4%estimated ± 1.2 pp, medium confidence
Claude Sonnet 4.527.7%76.4%estimated ± 1.2 pp, medium confidence
Claude Sonnet 4.636.9%78.4%estimated ± 1.2 pp, medium confidence
Gemini 3 Flash11.4%72.8%estimated ± 1.2 pp, medium confidence
Gemini 3 Pro11.4%72.8%estimated ± 1.2 pp, medium confidence
GPT-5.1-Codex26.2%76.1%estimated ± 1.2 pp, medium confidence
GPT-5.234.3%77.8%estimated ± 1.2 pp, medium confidence
GPT-5.2-Codex26.0%76.0%estimated ± 1.2 pp, medium confidence
GPT-5.3 Codex33.7%77.7%estimated ± 1.2 pp, medium confidence
GPT-5.438.9%78.8%estimated ± 1.2 pp, medium confidence
GPT-5.542.7%79.7%estimated ± 1.2 pp, medium confidence
GPT-5 (high)8.5%72.2%estimated ± 1.2 pp, low confidence
Kimi K352.9%81.9%estimated ± 1.2 pp, medium confidence
MiMo-V2.6-Flash61.2%83.7%estimated ± 1.2 pp, medium confidence
MiMo-V2.6-Pro62.0%83.9%estimated ± 1.2 pp, low confidence
Muse Spark 1.154.7%82.3%estimated ± 1.2 pp, medium confidence
Muse Spark 1.364.9%84.5%estimated ± 1.2 pp, low confidence
Qwen3.5 Plus18.5%74.4%estimated ± 1.2 pp, medium confidence
Qwen3.8-27B33.4%77.6%estimated ± 1.2 pp, medium confidence
Qwen3.8-Flash-Next55.7%82.5%estimated ± 1.2 pp, medium confidence
Step 5 Preview59.0%83.2%estimated ± 1.2 pp, medium confidence