benchgap
Calibration

Gert Labs → Claw-Eval

Claw-Eval is estimated from Gert Labs with a Michaelis–Menten curve fitted on 22 models measured on both: y = 0.8794·x / (0.23590 + x), R² = 0.36, cross-validated error 6.8 pp. It is used for 18 estimates.

Estimated modelGert LabsClaw-EvalSource
Gemini 2.5 Pro42.0%56.3%estimated ± 6.8 pp, low confidence
Gemini 3.1 Flash-Lite38.5%54.5%estimated ± 6.8 pp, low confidence
Gemma 4 31B35.3%52.7%estimated ± 6.8 pp, low confidence
GLM-4.740.0%55.3%estimated ± 6.8 pp, low confidence
GPT-4.125.7%45.8%estimated ± 6.8 pp, low confidence
GPT-5.141.2%55.9%estimated ± 6.8 pp, low confidence
GPT-OSS 120B29.6%48.9%estimated ± 6.8 pp, low confidence
Grok 442.3%56.5%estimated ± 6.8 pp, low confidence
Grok 4.1 Fast47.3%58.7%estimated ± 6.8 pp, low confidence
Grok 4.2038.4%54.5%estimated ± 6.8 pp, low confidence
Grok Build 0.149.2%59.4%estimated ± 6.8 pp, low confidence
Hy3 Preview36.9%53.6%estimated ± 6.8 pp, low confidence
Kimi K2.5 (Reasoning)32.6%51.0%estimated ± 6.8 pp, low confidence
Mistral Medium 3.5 128B39.1%54.8%estimated ± 6.8 pp, low confidence
Qwen3.5-27B39.4%55.0%estimated ± 6.8 pp, low confidence
Qwen3.5-35B-A3B29.0%48.5%estimated ± 6.8 pp, low confidence
Qwen3 Max43.7%57.1%estimated ± 6.8 pp, low confidence
Trinity-Large-Thinking32.6%51.0%estimated ± 6.8 pp, low confidence