Calibration
React Native Evals → Vals LiveCodeBench
Vals LiveCodeBench is estimated from React Native Evals with a Michaelis–Menten curve fitted on 5 models measured on both: y = 1.2755·x / (0.40962 + x), R² = 0.34, cross-validated error 3.8 pp. It is used for 9 estimates.
| Estimated model | React Native Evals | Vals LiveCodeBench | Source |
|---|---|---|---|
| Composer 2 | 96.1% | 89.4% | estimated ± 3.8 pp, low confidence |
| Composer 2 Fast | 94.9% | 89.1% | estimated ± 3.8 pp, low confidence |
| DeepSeek V3.2 | 71.5% | 81.1% | estimated ± 3.8 pp, low confidence |
| Gemma 4 31B | 75.2% | 82.6% | estimated ± 3.8 pp, low confidence |
| GLM-5 | 74.8% | 82.4% | estimated ± 3.8 pp, low confidence |
| GPT-5.4 | 85.3% | 86.2% | estimated ± 3.8 pp, low confidence |
| GPT-OSS 120B | 71.6% | 81.1% | estimated ± 3.8 pp, low confidence |
| GPT-OSS 20B | 71.0% | 80.9% | estimated ± 3.8 pp, low confidence |
| Grok 4 | 72.6% | 81.5% | estimated ± 3.8 pp, low confidence |