Calibration
Vals SWE-bench → React Native Evals
React Native Evals is estimated from Vals SWE-bench with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 1.21927·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.83, cross-validated error 2.4 pp. It is used for 32 estimates.
| Estimated model | Vals SWE-bench | React Native Evals | Source |
|---|---|---|---|
| GLM-5.3 | 95.4% | 100.0% | estimated ± 2.4 pp, low confidence |
| GLM-5.3-Flash | 92.0% | 100.0% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Luna | 93.0% | 100.0% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Sol | 96.2% | 100.0% | estimated ± 2.4 pp, low confidence |
| GPT-5.6 Terra | 95.4% | 100.0% | estimated ± 2.4 pp, low confidence |
| Grok 4.6 | 95.6% | 100.0% | estimated ± 2.4 pp, low confidence |
| Kimi K3 | 93.4% | 100.0% | estimated ± 2.4 pp, low confidence |
| Grok 4.5 | 86.6% | 93.1% | estimated ± 2.4 pp, low confidence |
| Muse Spark 1.2 | 86.6% | 93.1% | estimated ± 2.4 pp, low confidence |
| Qwen3.8-27B | 86.0% | 92.0% | estimated ± 2.4 pp, low confidence |
| Qwen3.8 Max | 85.6% | 91.2% | estimated ± 2.4 pp, low confidence |
| GLM-5.2 | 82.8% | 86.1% | estimated ± 2.4 pp, low confidence |
| Muse Spark 1.1 | 82.0% | 84.7% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.7 Flash | 80.8% | 82.6% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.8 Flash | 80.0% | 81.3% | estimated ± 2.4 pp, medium confidence |
| Composer 2.5 | 79.6% | 80.6% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.6 Flash | 79.6% | 80.6% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.5 Flash | 78.8% | 79.3% | estimated ± 2.4 pp, medium confidence |
| Kimi K2.7 Code | 78.2% | 78.3% | estimated ± 2.4 pp, medium confidence |
| GLM-5.1 | 76.4% | 75.4% | estimated ± 2.4 pp, medium confidence |
| Gemini 3.5 Flash-Lite | 75.0% | 73.2% | estimated ± 2.4 pp, medium confidence |
| Gemini 3 Flash | 75.0% | 73.2% | estimated ± 2.4 pp, medium confidence |
| MiMo-V2.5-Pro | 74.0% | 71.6% | estimated ± 2.4 pp, medium confidence |
| GPT-5.4 mini | 73.0% | 70.1% | estimated ± 2.4 pp, low confidence |
| Qwen 3.6 Max (preview) | 72.8% | 69.8% | estimated ± 2.4 pp, low confidence |
| GPT-5.2-Codex | 72.4% | 69.2% | estimated ± 2.4 pp, low confidence |
| Grok 4.3 | 71.4% | 67.7% | estimated ± 2.4 pp, low confidence |
| MiMo-V2.5 | 71.0% | 67.1% | estimated ± 2.4 pp, low confidence |
| GPT-5.4 nano | 69.8% | 65.4% | estimated ± 2.4 pp, low confidence |
| Ling 3.0 Flash | 65.2% | 59.0% | estimated ± 2.4 pp, low confidence |
| Qwen3.5 Flash | 64.4% | 57.9% | estimated ± 2.4 pp, low confidence |
| Gemini 3.1 Flash-Lite | 62.8% | 55.8% | estimated ± 2.4 pp, low confidence |