benchgap
Calibration

Claw-Eval → BFCL v4

BFCL v4 is estimated from Claw-Eval with a linear curve fitted on 5 models measured on both: y = 2.8304·x + -1.1404, R² = 0.60, cross-validated error 8.6 pp. It is used for 20 estimates.

Estimated modelClaw-EvalBFCL v4Source
Claude Opus 4.559.6%54.6%estimated ± 8.6 pp, low confidence
Claude Opus 4.670.4%85.2%estimated ± 8.6 pp, low confidence
Claude Sonnet 4.667.8%77.9%estimated ± 8.6 pp, low confidence
DeepSeek V3.240.2%0.0%estimated ± 8.6 pp, low confidence
dots3-note Preview73.4%93.7%estimated ± 8.6 pp, low confidence
Gemini 3 Flash49.2%25.2%estimated ± 8.6 pp, low confidence
GLM-557.7%49.3%estimated ± 8.6 pp, low confidence
GLM-5-Turbo55.8%43.9%estimated ± 8.6 pp, low confidence
GLM-5V-Turbo53.8%38.2%estimated ± 8.6 pp, low confidence
K-EXAONE 2.077.7%100.0%estimated ± 8.6 pp, low confidence
MiMo-V2.562.3%62.3%estimated ± 8.6 pp, low confidence
MiMo-V2-Omni45.2%13.9%estimated ± 8.6 pp, low confidence
MiMo-V2-Pro57.8%49.6%estimated ± 8.6 pp, low confidence
Ornith-1.0-35B69.8%83.5%estimated ± 8.6 pp, low confidence
Ornith-1.0-397B77.1%100.0%estimated ± 8.6 pp, low confidence
Ornith-1.0-9B63.1%64.6%estimated ± 8.6 pp, low confidence
Ornith-1.5-35B-A3B72.5%91.2%estimated ± 8.6 pp, low confidence
Ornith-1.5-397B81.4%100.0%estimated ± 8.6 pp, low confidence
Ornith-1.5-9B66.5%74.2%estimated ± 8.6 pp, low confidence
Qwen3.5 397B56.8%46.7%estimated ± 8.6 pp, low confidence