benchgap
Calibration

Claw-Eval → VITA-Bench

VITA-Bench is estimated from Claw-Eval with a inverse Michaelis–Menten curve fitted on 7 models measured on both: y = 0.87521·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.38, cross-validated error 10.7 pp. It is used for 18 estimates.

Estimated modelClaw-EvalVITA-BenchSource
Claude Sonnet 4.667.8%44.9%estimated ± 10.7 pp, low confidence
Gemini 3.1 Pro57.8%35.6%estimated ± 10.7 pp, low confidence
Gemini 3 Flash49.2%28.6%estimated ± 10.7 pp, low confidence
GLM-5-Turbo55.8%33.9%estimated ± 10.7 pp, low confidence
GLM-5V-Turbo53.8%32.2%estimated ± 10.7 pp, low confidence
K-EXAONE 2.077.7%55.6%estimated ± 10.7 pp, low confidence
LFM2.5-2.6B62.8%40.1%estimated ± 10.7 pp, low confidence
LLaDA2.2-mini57.2%35.0%estimated ± 10.7 pp, low confidence
MiMo-V2.562.3%39.6%estimated ± 10.7 pp, low confidence
MiMo-V2.5-Pro63.8%41.0%estimated ± 10.7 pp, low confidence
MiMo-V2-Omni45.2%25.6%estimated ± 10.7 pp, low confidence
MiMo-V2-Pro57.8%35.6%estimated ± 10.7 pp, low confidence
MiniMax M2.748.7%28.2%estimated ± 10.7 pp, low confidence
Muse Spark63.8%41.0%estimated ± 10.7 pp, low confidence
Nemotron 3 Super 100B5.5%2.5%estimated ± 10.7 pp, low confidence
Ornith-1.0-35B69.8%46.9%estimated ± 10.7 pp, low confidence
Ornith-1.0-397B77.1%54.9%estimated ± 10.7 pp, low confidence
Ornith-1.0-9B63.1%40.3%estimated ± 10.7 pp, low confidence