benchgap
Calibration

Gert Labs → ResearchClawBench

ResearchClawBench is estimated from Gert Labs with a linear curve fitted on 16 models measured on both: y = 0.1581·x + 0.0796, R² = 0.41, cross-validated error 2.2 pp. It is used for 13 estimates.

Estimated modelGert LabsResearchClawBenchSource
Claude 4 Sonnet39.7%14.2%estimated ± 2.2 pp, medium confidence
Claude Sonnet 4.548.5%15.6%estimated ± 2.2 pp, medium confidence
Gemini 3.1 Flash-Lite38.5%14.0%estimated ± 2.2 pp, medium confidence
Gemini 3 Pro63.2%18.0%estimated ± 2.2 pp, medium confidence
GPT-4.125.7%12.0%estimated ± 2.2 pp, low confidence
GPT-5.1-Codex49.7%15.8%estimated ± 2.2 pp, medium confidence
GPT-5.2-Codex51.8%16.1%estimated ± 2.2 pp, medium confidence
GPT-5.3 Codex57.5%17.0%estimated ± 2.2 pp, medium confidence
Grok 442.3%14.7%estimated ± 2.2 pp, medium confidence
Grok 4.1 Fast47.3%15.4%estimated ± 2.2 pp, medium confidence
Grok 4.2038.4%14.0%estimated ± 2.2 pp, medium confidence
Grok Build 0.149.2%15.7%estimated ± 2.2 pp, medium confidence
Qwen3 Max43.7%14.9%estimated ± 2.2 pp, medium confidence