benchgap
Calibration

JobBench → DeepSearchQA

DeepSearchQA is estimated from JobBench with a linear curve fitted on 7 models measured on both: y = 0.3350·x + 0.6954, R² = 0.40, cross-validated error 10.9 pp. It is used for 9 estimates.

Estimated modelJobBenchDeepSearchQASource
Claude 4.1 Opus21.9%76.9%estimated ± 10.9 pp, low confidence
Claude 4 Sonnet18.4%75.7%estimated ± 10.9 pp, low confidence
Claude Haiku 4.516.0%74.9%estimated ± 10.9 pp, low confidence
Gemini 3 Flash11.4%73.4%estimated ± 10.9 pp, low confidence
Gemini 3 Pro11.4%73.4%estimated ± 10.9 pp, low confidence
GPT-5.1-Codex26.2%78.3%estimated ± 10.9 pp, low confidence
GPT-5.2-Codex26.0%78.3%estimated ± 10.9 pp, low confidence
Hy4 preview61.7%90.2%estimated ± 10.9 pp, low confidence
Qwen3.5 Plus18.5%75.7%estimated ± 10.9 pp, low confidence