Calibration
JobBench → DeepSearchQA
DeepSearchQA is estimated from JobBench with a linear curve fitted on 7 models measured on both: y = 0.3350·x + 0.6954, R² = 0.40, cross-validated error 10.9 pp. It is used for 9 estimates.
| Estimated model | JobBench | DeepSearchQA | Source |
|---|---|---|---|
| Claude 4.1 Opus | 21.9% | 76.9% | estimated ± 10.9 pp, low confidence |
| Claude 4 Sonnet | 18.4% | 75.7% | estimated ± 10.9 pp, low confidence |
| Claude Haiku 4.5 | 16.0% | 74.9% | estimated ± 10.9 pp, low confidence |
| Gemini 3 Flash | 11.4% | 73.4% | estimated ± 10.9 pp, low confidence |
| Gemini 3 Pro | 11.4% | 73.4% | estimated ± 10.9 pp, low confidence |
| GPT-5.1-Codex | 26.2% | 78.3% | estimated ± 10.9 pp, low confidence |
| GPT-5.2-Codex | 26.0% | 78.3% | estimated ± 10.9 pp, low confidence |
| Hy4 preview | 61.7% | 90.2% | estimated ± 10.9 pp, low confidence |
| Qwen3.5 Plus | 18.5% | 75.7% | estimated ± 10.9 pp, low confidence |