Calibration
Gert Labs → Claw-Eval
Claw-Eval is estimated from Gert Labs with a Michaelis–Menten curve fitted on 22 models measured on both: y = 0.8794·x / (0.23590 + x), R² = 0.36, cross-validated error 6.8 pp. It is used for 18 estimates.
| Estimated model | Gert Labs | Claw-Eval | Source |
|---|---|---|---|
| Gemini 2.5 Pro | 42.0% | 56.3% | estimated ± 6.8 pp, low confidence |
| Gemini 3.1 Flash-Lite | 38.5% | 54.5% | estimated ± 6.8 pp, low confidence |
| Gemma 4 31B | 35.3% | 52.7% | estimated ± 6.8 pp, low confidence |
| GLM-4.7 | 40.0% | 55.3% | estimated ± 6.8 pp, low confidence |
| GPT-4.1 | 25.7% | 45.8% | estimated ± 6.8 pp, low confidence |
| GPT-5.1 | 41.2% | 55.9% | estimated ± 6.8 pp, low confidence |
| GPT-OSS 120B | 29.6% | 48.9% | estimated ± 6.8 pp, low confidence |
| Grok 4 | 42.3% | 56.5% | estimated ± 6.8 pp, low confidence |
| Grok 4.1 Fast | 47.3% | 58.7% | estimated ± 6.8 pp, low confidence |
| Grok 4.20 | 38.4% | 54.5% | estimated ± 6.8 pp, low confidence |
| Grok Build 0.1 | 49.2% | 59.4% | estimated ± 6.8 pp, low confidence |
| Hy3 Preview | 36.9% | 53.6% | estimated ± 6.8 pp, low confidence |
| Kimi K2.5 (Reasoning) | 32.6% | 51.0% | estimated ± 6.8 pp, low confidence |
| Mistral Medium 3.5 128B | 39.1% | 54.8% | estimated ± 6.8 pp, low confidence |
| Qwen3.5-27B | 39.4% | 55.0% | estimated ± 6.8 pp, low confidence |
| Qwen3.5-35B-A3B | 29.0% | 48.5% | estimated ± 6.8 pp, low confidence |
| Qwen3 Max | 43.7% | 57.1% | estimated ± 6.8 pp, low confidence |
| Trinity-Large-Thinking | 32.6% | 51.0% | estimated ± 6.8 pp, low confidence |