benchgap
Calibration

BrowseComp → VITA-Bench

VITA-Bench is estimated from BrowseComp with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.4811 − 0.0000)·x^6.00 / (0.52273^6.00 + x^6.00), R² = 0.74, cross-validated error 9.9 pp. It is used for 19 estimates.

Estimated modelBrowseCompVITA-BenchSource
Atria Dawn Preview92.5%46.6%estimated ± 9.9 pp, low confidence
Claude Mythos 588.0%46.1%estimated ± 9.9 pp, low confidence
Claude Opus 4.683.7%45.4%estimated ± 9.9 pp, low confidence
Claude Sonnet 584.7%45.6%estimated ± 9.9 pp, low confidence
dots3-note Preview83.3%45.3%estimated ± 9.9 pp, low confidence
GPT-5.265.8%38.4%estimated ± 9.9 pp, low confidence
GPT-5.4 Pro89.3%46.2%estimated ± 9.9 pp, low confidence
GPT-5.5 Pro90.1%46.3%estimated ± 9.9 pp, low confidence
GPT-5.6 Luna83.3%45.3%estimated ± 9.9 pp, low confidence
GPT-5.6 Sol92.2%46.6%estimated ± 9.9 pp, low confidence
GPT-5.6 Terra87.5%46.0%estimated ± 9.9 pp, low confidence
GPT-6 Astra91.5%46.5%estimated ± 9.9 pp, low confidence
Kimi K2.5 (Reasoning)60.6%34.1%estimated ± 9.9 pp, low confidence
Nemotron 3.5 Lightning 30B A3B NVFP436.8%5.2%estimated ± 9.9 pp, low confidence
Nemotron 3 Ultra44.4%13.1%estimated ± 9.9 pp, low confidence
Qwen3.5-122B-A10B63.8%36.9%estimated ± 9.9 pp, low confidence
Qwen3.5-27B61.0%34.5%estimated ± 9.9 pp, low confidence
Qwen3.5-35B-A3B61.0%34.5%estimated ± 9.9 pp, low confidence
Step 3.7 Flash75.8%43.4%estimated ± 9.9 pp, low confidence