Calibration
Vals SWE-bench → React Native Evals
React Native Evals is estimated from Vals SWE-bench with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 1.21927·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.83, cross-validated error 2.4 pp. It is used for 55 estimates.
| Estimated model | Vals SWE-bench | React Native Evals | Source |
|---|---|---|---|
| Claude Fable 5 | 95.0% | 100.0% | estimated ± 2.4 pp, low confidence |
| Claude Haiku 4.5 | 66.6% | 60.9% | estimated ± 2.4 pp, low confidence |
| Claude Opus 4.8 | 88.6% | 97.0% | estimated ± 2.4 pp, low confidence |
| Claude Opus 5 | 97.0% | 100.0% | estimated ± 2.4 pp, low confidence |
| Claude Sonnet 5 | 79.6% | 80.6% | estimated ± 2.4 pp, medium confidence |
| Composer 2.5 | 79.6% | 80.6% | estimated ± 2.4 pp, medium confidence |
| DeepSeek V4 Flash 0731 | 88.8% | 97.4% | estimated ± 2.4 pp, low confidence |
| DeepSeek V4 Pro 0813 | 96.4% | 100.0% | estimated ± 2.4 pp, low confidence |
| Gemini 2.5 Pro | 54.4% | 45.6% | estimated ± 2.4 pp, low confidence |
| Gemini 3.1 Flash-Lite | 62.8% | 55.8% | estimated ± 2.4 pp, low confidence |
| Gemini 3.5 Flash | 78.8% | 79.3% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.5 Flash-Lite | 75.0% | 73.2% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.6 Flash | 79.6% | 80.6% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.7 Flash | 80.8% | 82.6% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.8 Flash | 80.0% | 81.3% | estimated ± 2.4 pp, medium confidence |
| Gemini 3 Flash | 75.0% | 73.2% | estimated ± 2.4 pp, medium confidence |
| GLM-4.7 | 69.4% | 64.8% | estimated ± 2.4 pp, low confidence |
| GLM-5.1 | 76.4% | 75.4% | estimated ± 2.4 pp, medium confidence |
| GLM-5.2 | 82.8% | 86.1% | estimated ± 2.4 pp, low confidence |
| GLM-5.3 | 95.4% | 100.0% | estimated ± 2.4 pp, low confidence |
| GLM-5.3-Flash | 92.0% | 100.0% | estimated ± 2.4 pp, low confidence |
| GPT-5.2-Codex | 72.4% | 69.2% | estimated ± 2.4 pp, low confidence |
| GPT-5.3 Codex | 78.0% | 78.0% | estimated ± 2.4 pp, medium confidence |
| GPT-5.4 mini | 73.0% | 70.1% | estimated ± 2.4 pp, low confidence |
| GPT-5.4 nano | 69.8% | 65.4% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Luna | 93.0% | 100.0% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Sol | 96.2% | 100.0% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Terra | 95.4% | 100.0% | estimated ± 2.4 pp, low confidence |
| Grok 4.20 | 72.2% | 68.9% | estimated ± 2.4 pp, low confidence |
| Grok 4.3 | 71.4% | 67.7% | estimated ± 2.4 pp, low confidence |
| Grok 4.5 | 86.6% | 93.1% | estimated ± 2.4 pp, low confidence |
| Grok 4.6 | 95.6% | 100.0% | estimated ± 2.4 pp, low confidence |
| Inkling | 77.6% | 77.3% | estimated ± 2.4 pp, medium confidence |
| Inkling-Small | 82.2% | 85.1% | estimated ± 2.4 pp, medium confidence |
| Kimi K2.6 | 76.2% | 75.0% | estimated ± 2.4 pp, medium confidence |
| Kimi K2.7 Code | 78.2% | 78.3% | estimated ± 2.4 pp, medium confidence |
| Kimi K3 | 93.4% | 100.0% | estimated ± 2.4 pp, low confidence |
| Laguna M.1 | 57.6% | 49.3% | estimated ± 2.4 pp, low confidence |
| Laguna XS.2 | 55.2% | 46.5% | estimated ± 2.4 pp, low confidence |
| Ling 3.0 Flash | 65.2% | 59.0% | estimated ± 2.4 pp, low confidence |
| MiMo-V2.5 | 71.0% | 67.1% | estimated ± 2.4 pp, low confidence |
| MiMo-V2.5-Pro | 74.0% | 71.6% | estimated ± 2.4 pp, medium confidence |
| MiniMax M3 | 75.0% | 73.2% | estimated ± 2.4 pp, medium confidence |
| Mistral Medium 3.5 128B | 66.4% | 60.6% | estimated ± 2.4 pp, low confidence |
| Muse Spark | 74.4% | 72.2% | estimated ± 2.4 pp, medium confidence |
| Muse Spark 1.1 | 82.0% | 84.7% | estimated ± 2.4 pp, medium confidence |
| Muse Spark 1.2 | 86.6% | 93.1% | estimated ± 2.4 pp, low confidence |
| Nemotron 3 Ultra | 69.0% | 64.2% | estimated ± 2.4 pp, low confidence |
| Qwen3.5 Flash | 64.4% | 57.9% | estimated ± 2.4 pp, low confidence |
| Qwen3.6-27B | 70.0% | 65.7% | estimated ± 2.4 pp, low confidence |
| Qwen 3.6 Max (preview) | 72.8% | 69.8% | estimated ± 2.4 pp, low confidence |
| Qwen3.6 Plus | 73.4% | 70.7% | estimated ± 2.4 pp, low confidence |
| Qwen3.7 Max | 68.8% | 63.9% | estimated ± 2.4 pp, low confidence |
| Qwen3.8-27B | 86.0% | 92.0% | estimated ± 2.4 pp, low confidence |
| Qwen3.8 Max | 85.6% | 91.2% | estimated ± 2.4 pp, low confidence |