Calibration
Gert Labs → WideResearch
WideResearch is estimated from Gert Labs with a Hill curve fitted on 7 models measured on both: y = 0.0000 + (0.8099 − 0.0000)·x^6.00 / (0.34184^6.00 + x^6.00), R² = 0.67, cross-validated error 5.9 pp. It is used for 23 estimates.
| Estimated model | Gert Labs | WideResearch | Source |
|---|---|---|---|
| Claude Opus 4.7 | 65.6% | 79.4% | estimated ± 5.9 pp, low confidence |
| DeepSeek V3.2 | 29.6% | 23.9% | estimated ± 5.9 pp, low confidence |
| Gemini 2.5 Pro | 42.0% | 62.8% | estimated ± 5.9 pp, low confidence |
| Gemini 3.1 Flash-Lite | 38.5% | 54.2% | estimated ± 5.9 pp, low confidence |
| Gemini 3.1 Pro | 56.9% | 77.3% | estimated ± 5.9 pp, low confidence |
| Gemma 4 31B | 35.3% | 44.2% | estimated ± 5.9 pp, low confidence |
| GLM-5V-Turbo | 30.8% | 28.1% | estimated ± 5.9 pp, low confidence |
| GPT-4.1 | 25.7% | 12.3% | estimated ± 5.9 pp, low confidence |
| GPT-5.1 | 41.2% | 61.2% | estimated ± 5.9 pp, low confidence |
| GPT-OSS 120B | 29.6% | 24.0% | estimated ± 5.9 pp, low confidence |
| Grok 4 | 42.3% | 63.4% | estimated ± 5.9 pp, low confidence |
| Grok 4.1 Fast | 47.3% | 70.9% | estimated ± 5.9 pp, low confidence |
| Grok 4.20 | 38.4% | 54.0% | estimated ± 5.9 pp, low confidence |
| Grok 4.3 | 43.9% | 66.2% | estimated ± 5.9 pp, low confidence |
| Grok Build 0.1 | 49.2% | 72.8% | estimated ± 5.9 pp, low confidence |
| Hy3 Preview | 36.9% | 49.7% | estimated ± 5.9 pp, low confidence |
| MiMo-V2.5 | 46.9% | 70.4% | estimated ± 5.9 pp, low confidence |
| MiMo-V2.5-Pro | 62.7% | 78.9% | estimated ± 5.9 pp, low confidence |
| MiMo-V2-Pro | 36.7% | 48.9% | estimated ± 5.9 pp, low confidence |
| Mistral Medium 3.5 128B | 39.1% | 56.0% | estimated ± 5.9 pp, low confidence |
| Qwen3.6-27B | 54.8% | 76.5% | estimated ± 5.9 pp, low confidence |
| Qwen3 Max | 43.7% | 66.0% | estimated ± 5.9 pp, low confidence |
| Trinity-Large-Thinking | 32.6% | 34.6% | estimated ± 5.9 pp, low confidence |