benchgap
Calibration

Gert Labs → BrowseComp

BrowseComp is estimated from Gert Labs with a offset logistic curve fitted on 14 models measured on both: y = 0.5863 + (0.8111 − 0.5863) / (1 + exp(−53.46·(x − 0.4924))), R² = 0.84, cross-validated error 5.7 pp. It is used for 24 estimates.

Estimated modelGert LabsBrowseCompSource
Claude 4 Sonnet39.7%58.8%estimated ± 5.7 pp, medium confidence
DeepSeek V3.229.6%58.6%estimated ± 5.7 pp, medium confidence
Gemini 2.5 Pro42.0%59.1%estimated ± 5.7 pp, medium confidence
Gemini 3.1 Flash-Lite38.5%58.7%estimated ± 5.7 pp, medium confidence
Gemini 3 Flash56.6%80.7%estimated ± 5.7 pp, medium confidence
Gemini 3 Pro63.2%81.1%estimated ± 5.7 pp, medium confidence
Gemma 4 31B35.3%58.6%estimated ± 5.7 pp, medium confidence
GLM-5V-Turbo30.8%58.6%estimated ± 5.7 pp, medium confidence
GPT-4.125.7%58.6%estimated ± 5.7 pp, low confidence
GPT-5.141.2%58.9%estimated ± 5.7 pp, medium confidence
GPT-5.1-Codex49.7%71.2%estimated ± 5.7 pp, medium confidence
GPT-5.2-Codex51.8%76.5%estimated ± 5.7 pp, medium confidence
Grok 442.3%59.2%estimated ± 5.7 pp, medium confidence
Grok 4.1 Fast47.3%64.6%estimated ± 5.7 pp, medium confidence
Grok 4.2038.4%58.7%estimated ± 5.7 pp, medium confidence
Grok Build 0.149.2%69.6%estimated ± 5.7 pp, medium confidence
Hy3 Preview36.9%58.7%estimated ± 5.7 pp, medium confidence
MiMo-V2.546.9%63.6%estimated ± 5.7 pp, medium confidence
MiMo-V2-Pro36.7%58.7%estimated ± 5.7 pp, medium confidence
Mistral Medium 3.5 128B39.1%58.7%estimated ± 5.7 pp, medium confidence
Qwen3.6-27B54.8%80.0%estimated ± 5.7 pp, medium confidence
Qwen3.7 Max64.3%81.1%estimated ± 5.7 pp, medium confidence
Qwen3 Max43.7%59.8%estimated ± 5.7 pp, medium confidence
Trinity-Large-Thinking32.6%58.6%estimated ± 5.7 pp, medium confidence