benchgap
Calibration

GDPval-AA → BrowseComp

BrowseComp is estimated from GDPval-AA with a Michaelis–Menten curve fitted on 28 models measured on both: y = 1.1061·x / (0.15563 + x), R² = 0.59, cross-validated error 10.1 pp. It is used for 23 estimates.

Estimated modelGDPval-AABrowseCompSource
A.X K223.2%66.2%estimated ± 10.1 pp, low confidence
DeepSeek V3 03240.0%0.0%estimated ± 10.1 pp, low confidence
Gemma 4 12B0.0%0.0%estimated ± 10.1 pp, low confidence
Gemma 4 26B A4B3.4%19.8%estimated ± 10.1 pp, low confidence
Gemma 4 E2B0.0%0.0%estimated ± 10.1 pp, low confidence
Gemma 4 E4B0.0%0.0%estimated ± 10.1 pp, low confidence
GPT-4.1 mini0.0%0.0%estimated ± 10.1 pp, low confidence
GPT-4.1 nano0.0%0.0%estimated ± 10.1 pp, low confidence
GPT-4o0.0%0.0%estimated ± 10.1 pp, low confidence
GPT-4o mini0.0%0.0%estimated ± 10.1 pp, low confidence
Granite 4.2 30B3.9%22.2%estimated ± 10.1 pp, low confidence
Granite 4.2 3B0.0%0.0%estimated ± 10.1 pp, low confidence
K-Exaone0.0%0.0%estimated ± 10.1 pp, low confidence
Ling 2.6 Flash0.0%0.0%estimated ± 10.1 pp, low confidence
Ling 3.0 Flash VL33.2%75.3%estimated ± 10.1 pp, low confidence
Ling 3.0 Tiny3.2%18.9%estimated ± 10.1 pp, low confidence
MiMo-V2-Flash6.6%32.9%estimated ± 10.1 pp, low confidence
MiniCPM5-2B10.5%44.6%estimated ± 10.1 pp, low confidence
Mistral Large 446.2%82.7%estimated ± 10.1 pp, low confidence
Nemotron 3 Nano Omni 30B A3B0.0%0.0%estimated ± 10.1 pp, low confidence
North Mini Code0.0%0.0%estimated ± 10.1 pp, low confidence
Solar Pro 30.0%0.0%estimated ± 10.1 pp, low confidence
Ultravox v0.6 Llama 3.3 70B0.0%0.0%estimated ± 10.1 pp, low confidence