Calibration
BrowseComp → ResearchClawBench
ResearchClawBench is estimated from BrowseComp with a Michaelis–Menten + offset curve fitted on 9 models measured on both: y = 0.0000 + 0.6410·x / (2.03915 + x), R² = 0.47, cross-validated error 2.2 pp. It is used for 11 estimates.
| Estimated model | BrowseComp | ResearchClawBench | Source |
|---|---|---|---|
| Agents-A1 | 75.5% | 17.3% | estimated ± 2.2 pp, medium confidence |
| Agents-A1-4B | 66.8% | 15.8% | estimated ± 2.2 pp, medium confidence |
| Atria Dawn Preview | 92.5% | 20.0% | estimated ± 2.2 pp, low confidence |
| Claude Mythos 5 | 88.0% | 19.3% | estimated ± 2.2 pp, low confidence |
| GLM-4.7 | 52.0% | 13.0% | estimated ± 2.2 pp, low confidence |
| GPT-5.2 | 65.8% | 15.6% | estimated ± 2.2 pp, medium confidence |
| GPT-5.4 Pro | 89.3% | 19.5% | estimated ± 2.2 pp, low confidence |
| GPT-5.5 Pro | 90.1% | 19.6% | estimated ± 2.2 pp, low confidence |
| Kimi K2.5 (Reasoning) | 60.6% | 14.7% | estimated ± 2.2 pp, medium confidence |
| Qwen3.5-27B | 61.0% | 14.8% | estimated ± 2.2 pp, medium confidence |
| Qwen3.5-35B-A3B | 61.0% | 14.8% | estimated ± 2.2 pp, medium confidence |