benchgap
Calibration

React Native Evals → Vibe Code Bench

Vibe Code Bench is estimated from React Native Evals with a offset logistic curve fitted on 7 models measured on both: y = 0.2767 + (0.6655 − 0.2767) / (1 + exp(−166.46·(x − 0.8029))), R² = 0.94, cross-validated error 6.5 pp. It is used for 9 estimates.

Estimated modelReact Native EvalsVibe Code BenchSource
Composer 296.1%66.5%estimated ± 6.5 pp, low confidence
Composer 2 Fast94.9%66.5%estimated ± 6.5 pp, low confidence
DeepSeek V3.271.5%27.7%estimated ± 6.5 pp, low confidence
Gemma 4 31B75.2%27.7%estimated ± 6.5 pp, low confidence
GLM-574.8%27.7%estimated ± 6.5 pp, low confidence
GPT-OSS 120B71.6%27.7%estimated ± 6.5 pp, low confidence
GPT-OSS 20B71.0%27.7%estimated ± 6.5 pp, low confidence
Grok 472.6%27.7%estimated ± 6.5 pp, low confidence
Kimi K2.577.2%27.9%estimated ± 6.5 pp, low confidence