Calibration
BrowseComp → VITA-Bench
VITA-Bench is estimated from BrowseComp with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.4811 − 0.0000)·x^6.00 / (0.52273^6.00 + x^6.00), R² = 0.74, cross-validated error 9.9 pp. It is used for 19 estimates.
| Estimated model | BrowseComp | VITA-Bench | Source |
|---|---|---|---|
| Atria Dawn Preview | 92.5% | 46.6% | estimated ± 9.9 pp, low confidence |
| Claude Mythos 5 | 88.0% | 46.1% | estimated ± 9.9 pp, low confidence |
| Claude Opus 4.6 | 83.7% | 45.4% | estimated ± 9.9 pp, low confidence |
| Claude Sonnet 5 | 84.7% | 45.6% | estimated ± 9.9 pp, low confidence |
| dots3-note Preview | 83.3% | 45.3% | estimated ± 9.9 pp, low confidence |
| GPT-5.2 | 65.8% | 38.4% | estimated ± 9.9 pp, low confidence |
| GPT-5.4 Pro | 89.3% | 46.2% | estimated ± 9.9 pp, low confidence |
| GPT-5.5 Pro | 90.1% | 46.3% | estimated ± 9.9 pp, low confidence |
| GPT-5.6 Luna | 83.3% | 45.3% | estimated ± 9.9 pp, low confidence |
| GPT-5.6 Sol | 92.2% | 46.6% | estimated ± 9.9 pp, low confidence |
| GPT-5.6 Terra | 87.5% | 46.0% | estimated ± 9.9 pp, low confidence |
| GPT-6 Astra | 91.5% | 46.5% | estimated ± 9.9 pp, low confidence |
| Kimi K2.5 (Reasoning) | 60.6% | 34.1% | estimated ± 9.9 pp, low confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 36.8% | 5.2% | estimated ± 9.9 pp, low confidence |
| Nemotron 3 Ultra | 44.4% | 13.1% | estimated ± 9.9 pp, low confidence |
| Qwen3.5-122B-A10B | 63.8% | 36.9% | estimated ± 9.9 pp, low confidence |
| Qwen3.5-27B | 61.0% | 34.5% | estimated ± 9.9 pp, low confidence |
| Qwen3.5-35B-A3B | 61.0% | 34.5% | estimated ± 9.9 pp, low confidence |
| Step 3.7 Flash | 75.8% | 43.4% | estimated ± 9.9 pp, low confidence |