benchgap
Calibration

JobBench → Claw-Eval

Claw-Eval is estimated from JobBench with a offset logistic curve fitted on 6 models measured on both: y = 0.5075 + (0.6617 − 0.5075) / (1 + exp(−200.00·(x − 0.3215))), R² = 0.83, cross-validated error 5.6 pp. It is used for 22 estimates.

Estimated modelJobBenchClaw-EvalSource
Claude 4.1 Opus21.9%50.7%estimated ± 5.6 pp, low confidence
Claude 4 Sonnet18.4%50.7%estimated ± 5.6 pp, low confidence
Claude Haiku 4.516.0%50.7%estimated ± 5.6 pp, low confidence
Claude Opus 4.7 (Adaptive)45.9%66.2%estimated ± 5.6 pp, low confidence
Claude Sonnet 4.527.7%50.8%estimated ± 5.6 pp, low confidence
Gemini 3 Pro11.4%50.7%estimated ± 5.6 pp, low confidence
GPT-5.1-Codex26.2%50.7%estimated ± 5.6 pp, low confidence
GPT-5.234.3%66.0%estimated ± 5.6 pp, low confidence
GPT-5.2-Codex26.0%50.7%estimated ± 5.6 pp, low confidence
GPT-5.3 Codex33.7%65.5%estimated ± 5.6 pp, low confidence
GPT-5 (high)8.5%50.7%estimated ± 5.6 pp, low confidence
Hy4 preview61.7%66.2%estimated ± 5.6 pp, low confidence
Kimi K352.9%66.2%estimated ± 5.6 pp, low confidence
MiMo-V2.6-Flash61.2%66.2%estimated ± 5.6 pp, low confidence
MiMo-V2.6-Pro62.0%66.2%estimated ± 5.6 pp, low confidence
Muse Spark 1.154.7%66.2%estimated ± 5.6 pp, low confidence
Muse Spark 1.364.9%66.2%estimated ± 5.6 pp, low confidence
Qwen3.5 Plus18.5%50.7%estimated ± 5.6 pp, low confidence
Qwen3.8-27B33.4%65.0%estimated ± 5.6 pp, low confidence
Qwen3.8-Flash-Next55.7%66.2%estimated ± 5.6 pp, low confidence
Qwen3.8 Max53.4%66.2%estimated ± 5.6 pp, low confidence
Step 5 Preview59.0%66.2%estimated ± 5.6 pp, low confidence