benchgap
Calibration

Vibe Code Bench → React Native Evals

React Native Evals is estimated from Vibe Code Bench with a Michaelis–Menten curve fitted on 7 models measured on both: y = 0.9333·x / (0.07259 + x), R² = 0.86, cross-validated error 3.0 pp. It is used for 22 estimates.

Estimated modelVibe Code BenchReact Native EvalsSource
Claude Haiku 4.5 Thinking11.4%57.0%estimated ± 3.0 pp, low confidence
Claude Opus 4.5 Thinking20.6%69.0%estimated ± 3.0 pp, low confidence
Claude Opus 4.6 (Adaptive)53.5%82.2%estimated ± 3.0 pp, medium confidence
Claude Sonnet 4.5 Thinking22.6%70.7%estimated ± 3.0 pp, low confidence
DeepSeek V3.2 (Thinking)5.1%38.5%estimated ± 3.0 pp, low confidence
Gemini 3 Pro14.3%61.9%estimated ± 3.0 pp, low confidence
Gemini 4 Argon91.9%86.5%estimated ± 3.0 pp, low confidence
GLM-4.63.1%27.9%estimated ± 3.0 pp, low confidence
GLM-5 (Reasoning)23.4%71.2%estimated ± 3.0 pp, low confidence
GPT-5.124.6%72.1%estimated ± 3.0 pp, low confidence
GPT-5.1-Codex13.1%60.1%estimated ± 3.0 pp, low confidence
GPT-5.1-Codex-Max22.2%70.3%estimated ± 3.0 pp, low confidence
GPT-5.253.5%82.2%estimated ± 3.0 pp, medium confidence
GPT-5 (high)20.1%68.6%estimated ± 3.0 pp, low confidence
GPT-5 mini14.2%61.7%estimated ± 3.0 pp, low confidence
Grok 4.1 Fast (Reasoning)1.2%13.2%estimated ± 3.0 pp, low confidence
Grok 4 Fast (Reasoning)0.0%0.0%estimated ± 3.0 pp, low confidence
Kimi K2.5 (Reasoning)17.5%66.0%estimated ± 3.0 pp, low confidence
MiniMax M2.514.9%62.7%estimated ± 3.0 pp, low confidence
Mistral Large 478.4%85.4%estimated ± 3.0 pp, low confidence
Qwen3.5 Plus15.7%63.9%estimated ± 3.0 pp, low confidence
Qwen3 Max3.5%30.4%estimated ± 3.0 pp, low confidence