Calibration
BrowseComp → Claw-Eval
Claw-Eval is estimated from BrowseComp with a offset logistic curve fitted on 12 models measured on both: y = 0.6240 + (0.8170 − 0.6240) / (1 + exp(−200.00·(x − 0.8368))), R² = 0.51, cross-validated error 7.9 pp. It is used for 11 estimates.
| Estimated model | BrowseComp | Claw-Eval | Source |
|---|---|---|---|
| Claude Mythos 5 | 88.0% | 81.7% | estimated ± 7.9 pp, low confidence |
| Claude Sonnet 5 | 84.7% | 79.5% | estimated ± 7.9 pp, medium confidence |
| GPT-5.4 Pro | 89.3% | 81.7% | estimated ± 7.9 pp, low confidence |
| GPT-5.5 Pro | 90.1% | 81.7% | estimated ± 7.9 pp, low confidence |
| GPT-5.6 Luna | 83.3% | 68.5% | estimated ± 7.9 pp, medium confidence |
| GPT-5.6 Sol | 92.2% | 81.7% | estimated ± 7.9 pp, low confidence |
| GPT-5.6 Terra | 87.5% | 81.7% | estimated ± 7.9 pp, low confidence |
| GPT-6 Astra | 91.5% | 81.7% | estimated ± 7.9 pp, low confidence |
| Nemotron 3.5 Lightning 30B A3B NVFP4 | 36.8% | 62.4% | estimated ± 7.9 pp, low confidence |
| Nemotron 3 Ultra | 44.4% | 62.4% | estimated ± 7.9 pp, low confidence |
| Qwen3.5-122B-A10B | 63.8% | 62.4% | estimated ± 7.9 pp, medium confidence |