Calibration
JobBench → Claw-Eval
Claw-Eval is estimated from JobBench with a offset logistic curve fitted on 6 models measured on both: y = 0.5075 + (0.6617 − 0.5075) / (1 + exp(−200.00·(x − 0.3215))), R² = 0.83, cross-validated error 5.6 pp. It is used for 22 estimates.
| Estimated model | JobBench | Claw-Eval | Source |
|---|---|---|---|
| Claude 4.1 Opus | 21.9% | 50.7% | estimated ± 5.6 pp, low confidence |
| Claude 4 Sonnet | 18.4% | 50.7% | estimated ± 5.6 pp, low confidence |
| Claude Haiku 4.5 | 16.0% | 50.7% | estimated ± 5.6 pp, low confidence |
| Claude Opus 4.7 (Adaptive) | 45.9% | 66.2% | estimated ± 5.6 pp, low confidence |
| Claude Sonnet 4.5 | 27.7% | 50.8% | estimated ± 5.6 pp, low confidence |
| Gemini 3 Pro | 11.4% | 50.7% | estimated ± 5.6 pp, low confidence |
| GPT-5.1-Codex | 26.2% | 50.7% | estimated ± 5.6 pp, low confidence |
| GPT-5.2 | 34.3% | 66.0% | estimated ± 5.6 pp, low confidence |
| GPT-5.2-Codex | 26.0% | 50.7% | estimated ± 5.6 pp, low confidence |
| GPT-5.3 Codex | 33.7% | 65.5% | estimated ± 5.6 pp, low confidence |
| GPT-5 (high) | 8.5% | 50.7% | estimated ± 5.6 pp, low confidence |
| Hy4 preview | 61.7% | 66.2% | estimated ± 5.6 pp, low confidence |
| Kimi K3 | 52.9% | 66.2% | estimated ± 5.6 pp, low confidence |
| MiMo-V2.6-Flash | 61.2% | 66.2% | estimated ± 5.6 pp, low confidence |
| MiMo-V2.6-Pro | 62.0% | 66.2% | estimated ± 5.6 pp, low confidence |
| Muse Spark 1.1 | 54.7% | 66.2% | estimated ± 5.6 pp, low confidence |
| Muse Spark 1.3 | 64.9% | 66.2% | estimated ± 5.6 pp, low confidence |
| Qwen3.5 Plus | 18.5% | 50.7% | estimated ± 5.6 pp, low confidence |
| Qwen3.8-27B | 33.4% | 65.0% | estimated ± 5.6 pp, low confidence |
| Qwen3.8-Flash-Next | 55.7% | 66.2% | estimated ± 5.6 pp, low confidence |
| Qwen3.8 Max | 53.4% | 66.2% | estimated ± 5.6 pp, low confidence |
| Step 5 Preview | 59.0% | 66.2% | estimated ± 5.6 pp, low confidence |