benchgap
Calibration

Gert Labs → WideResearch

WideResearch is estimated from Gert Labs with a Hill curve fitted on 7 models measured on both: y = 0.0000 + (0.8099 − 0.0000)·x^6.00 / (0.34184^6.00 + x^6.00), R² = 0.67, cross-validated error 5.9 pp. It is used for 23 estimates.

Estimated modelGert LabsWideResearchSource
Claude Opus 4.765.6%79.4%estimated ± 5.9 pp, low confidence
DeepSeek V3.229.6%23.9%estimated ± 5.9 pp, low confidence
Gemini 2.5 Pro42.0%62.8%estimated ± 5.9 pp, low confidence
Gemini 3.1 Flash-Lite38.5%54.2%estimated ± 5.9 pp, low confidence
Gemini 3.1 Pro56.9%77.3%estimated ± 5.9 pp, low confidence
Gemma 4 31B35.3%44.2%estimated ± 5.9 pp, low confidence
GLM-5V-Turbo30.8%28.1%estimated ± 5.9 pp, low confidence
GPT-4.125.7%12.3%estimated ± 5.9 pp, low confidence
GPT-5.141.2%61.2%estimated ± 5.9 pp, low confidence
GPT-OSS 120B29.6%24.0%estimated ± 5.9 pp, low confidence
Grok 442.3%63.4%estimated ± 5.9 pp, low confidence
Grok 4.1 Fast47.3%70.9%estimated ± 5.9 pp, low confidence
Grok 4.2038.4%54.0%estimated ± 5.9 pp, low confidence
Grok 4.343.9%66.2%estimated ± 5.9 pp, low confidence
Grok Build 0.149.2%72.8%estimated ± 5.9 pp, low confidence
Hy3 Preview36.9%49.7%estimated ± 5.9 pp, low confidence
MiMo-V2.546.9%70.4%estimated ± 5.9 pp, low confidence
MiMo-V2.5-Pro62.7%78.9%estimated ± 5.9 pp, low confidence
MiMo-V2-Pro36.7%48.9%estimated ± 5.9 pp, low confidence
Mistral Medium 3.5 128B39.1%56.0%estimated ± 5.9 pp, low confidence
Qwen3.6-27B54.8%76.5%estimated ± 5.9 pp, low confidence
Qwen3 Max43.7%66.0%estimated ± 5.9 pp, low confidence
Trinity-Large-Thinking32.6%34.6%estimated ± 5.9 pp, low confidence