Calibration
BrowseComp → AA-AnalystAgent
AA-AnalystAgent is estimated from BrowseComp with a Hill curve fitted on 9 models measured on both: y = 0.0349 + (1.2000 − 0.0349)·x^6.00 / (1.00590^6.00 + x^6.00), R² = 0.63, cross-validated error 12.5 pp. It is used for 12 estimates.
| Estimated model | BrowseComp | AA-AnalystAgent | Source |
|---|---|---|---|
| Agents-A1-4B | 66.8% | 12.7% | estimated ± 12.5 pp, low confidence |
| Atria Dawn Preview | 92.5% | 47.4% | estimated ± 12.5 pp, low confidence |
| Claude Mythos 5 | 88.0% | 39.6% | estimated ± 12.5 pp, low confidence |
| Claude Opus 4.6 | 83.7% | 32.5% | estimated ± 12.5 pp, low confidence |
| GPT-5.2 | 65.8% | 12.0% | estimated ± 12.5 pp, low confidence |
| GPT-5.4 Pro | 89.3% | 41.8% | estimated ± 12.5 pp, low confidence |
| GPT-5.5 Pro | 90.1% | 43.2% | estimated ± 12.5 pp, low confidence |
| LongCat-Flash-Lite-Sparse | 48.6% | 5.0% | estimated ± 12.5 pp, low confidence |
| Qwen3.5-27B | 61.0% | 9.0% | estimated ± 12.5 pp, low confidence |
| Qwen3.5-35B-A3B | 61.0% | 9.0% | estimated ± 12.5 pp, low confidence |
| Qwen3.5 397B | 62.0% | 9.5% | estimated ± 12.5 pp, low confidence |
| Beam | 77.4% | 23.5% | estimated ± 12.5 pp, low confidence |