benchgap
Calibration

Vals SWE-bench → React Native Evals

React Native Evals is estimated from Vals SWE-bench with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 1.21927·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.83, cross-validated error 2.4 pp. It is used for 32 estimates.

Estimated modelVals SWE-benchReact Native EvalsSource
GLM-5.395.4%100.0%estimated ± 2.4 pp, low confidence
GLM-5.3-Flash92.0%100.0%estimated ± 2.4 pp, low confidence
GPT-5.6 Luna93.0%100.0%estimated ± 2.4 pp, low confidence
GPT-5.6 Sol96.2%100.0%estimated ± 2.4 pp, low confidence
GPT-5.6 Terra95.4%100.0%estimated ± 2.4 pp, low confidence
Grok 4.695.6%100.0%estimated ± 2.4 pp, low confidence
Kimi K393.4%100.0%estimated ± 2.4 pp, low confidence
Grok 4.586.6%93.1%estimated ± 2.4 pp, low confidence
Muse Spark 1.286.6%93.1%estimated ± 2.4 pp, low confidence
Qwen3.8-27B86.0%92.0%estimated ± 2.4 pp, low confidence
Qwen3.8 Max85.6%91.2%estimated ± 2.4 pp, low confidence
GLM-5.282.8%86.1%estimated ± 2.4 pp, low confidence
Muse Spark 1.182.0%84.7%estimated ± 2.4 pp, medium confidence
Gemini 3.7 Flash80.8%82.6%estimated ± 2.4 pp, medium confidence
Gemini 3.8 Flash80.0%81.3%estimated ± 2.4 pp, medium confidence
Composer 2.579.6%80.6%estimated ± 2.4 pp, medium confidence
Gemini 3.6 Flash79.6%80.6%estimated ± 2.4 pp, medium confidence
Gemini 3.5 Flash78.8%79.3%estimated ± 2.4 pp, medium confidence
Kimi K2.7 Code78.2%78.3%estimated ± 2.4 pp, medium confidence
GLM-5.176.4%75.4%estimated ± 2.4 pp, medium confidence
Gemini 3.5 Flash-Lite75.0%73.2%estimated ± 2.4 pp, medium confidence
Gemini 3 Flash75.0%73.2%estimated ± 2.4 pp, medium confidence
MiMo-V2.5-Pro74.0%71.6%estimated ± 2.4 pp, medium confidence
GPT-5.4 mini73.0%70.1%estimated ± 2.4 pp, low confidence
Qwen 3.6 Max (preview)72.8%69.8%estimated ± 2.4 pp, low confidence
GPT-5.2-Codex72.4%69.2%estimated ± 2.4 pp, low confidence
Grok 4.371.4%67.7%estimated ± 2.4 pp, low confidence
MiMo-V2.571.0%67.1%estimated ± 2.4 pp, low confidence
GPT-5.4 nano69.8%65.4%estimated ± 2.4 pp, low confidence
Ling 3.0 Flash65.2%59.0%estimated ± 2.4 pp, low confidence
Qwen3.5 Flash64.4%57.9%estimated ± 2.4 pp, low confidence
Gemini 3.1 Flash-Lite62.8%55.8%estimated ± 2.4 pp, low confidence