benchgap
Calibration

BFCL v4 → Claw-Eval

Claw-Eval is estimated from BFCL v4 with a Hill curve fitted on 5 models measured on both: y = 0.0000 + (0.6516 − 0.0000)·x^6.00 / (0.34036^6.00 + x^6.00), R² = 0.87, cross-validated error 2.3 pp. It is used for 17 estimates.

Estimated modelBFCL v4Claw-EvalSource
Atria Dawn Preview77.0%64.7%estimated ± 2.3 pp, low confidence
BTL-388.5%64.9%estimated ± 2.3 pp, low confidence
BTL-473.5%64.5%estimated ± 2.3 pp, medium confidence
Granite 4.2 30B61.4%63.3%estimated ± 2.3 pp, medium confidence
Granite 4.2 3B52.4%60.6%estimated ± 2.3 pp, medium confidence
Granite 4.2 8B52.4%60.6%estimated ± 2.3 pp, medium confidence
LFM2.5-230M21.0%3.4%estimated ± 2.3 pp, low confidence
LFM2.5-8B-A1B49.7%59.1%estimated ± 2.3 pp, medium confidence
LFM2.5-VL-3B32.5%28.1%estimated ± 2.3 pp, low confidence
LFM2.5-VL-450M21.1%3.5%estimated ± 2.3 pp, low confidence
Ling 3.0 Flash73.0%64.5%estimated ± 2.3 pp, medium confidence
Mellum2-12B-A2.5B-Instruct44.2%53.9%estimated ± 2.3 pp, low confidence
Mellum2-12B-A2.5B-Thinking45.6%55.6%estimated ± 2.3 pp, low confidence
MiniCPM5-1B25.2%9.1%estimated ± 2.3 pp, low confidence
MiniCPM5-2B66.6%64.0%estimated ± 2.3 pp, medium confidence
Pokee-Isaac 28B70.9%64.4%estimated ± 2.3 pp, medium confidence
ZAYA1-8B39.2%45.7%estimated ± 2.3 pp, low confidence