Calibration
Claw-Eval → JobBench
JobBench is estimated from Claw-Eval with a Hill curve fitted on 6 models measured on both: y = 0.0000 + (0.5463 − 0.0000)·x^6.00 / (0.58976^6.00 + x^6.00), R² = 0.77, cross-validated error 8.4 pp. It is used for 14 estimates.
| Estimated model | Claw-Eval | JobBench | Source |
|---|---|---|---|
| DeepSeek V3.2 | 40.2% | 5.0% | estimated ± 8.4 pp, low confidence |
| GLM-5-Turbo | 55.8% | 22.8% | estimated ± 8.4 pp, low confidence |
| GLM-5V-Turbo | 53.8% | 20.0% | estimated ± 8.4 pp, low confidence |
| K-EXAONE 2.0 | 77.7% | 45.9% | estimated ± 8.4 pp, low confidence |
| LFM2.5-2.6B | 62.8% | 32.5% | estimated ± 8.4 pp, low confidence |
| LLaDA2.2-mini | 57.2% | 24.8% | estimated ± 8.4 pp, low confidence |
| MiMo-V2.5 | 62.3% | 31.8% | estimated ± 8.4 pp, low confidence |
| MiMo-V2-Omni | 45.2% | 9.2% | estimated ± 8.4 pp, low confidence |
| MiMo-V2-Pro | 57.8% | 25.7% | estimated ± 8.4 pp, low confidence |
| Muse Spark | 63.8% | 33.6% | estimated ± 8.4 pp, low confidence |
| Ornith-1.0-35B | 69.8% | 40.1% | estimated ± 8.4 pp, low confidence |
| Ornith-1.0-397B | 77.1% | 45.5% | estimated ± 8.4 pp, low confidence |
| Ornith-1.0-9B | 63.1% | 32.8% | estimated ± 8.4 pp, low confidence |
| Qwen3.6-27B | 72.4% | 42.3% | estimated ± 8.4 pp, low confidence |