benchgap
Calibration

Vals SWE-bench → React Native Evals

React Native Evals is estimated from Vals SWE-bench with a inverse Michaelis–Menten curve fitted on 5 models measured on both: y = 1.21927·(x − 0.0000) / (0.0000 + 2.0000 − x), R² = 0.83, cross-validated error 2.4 pp. It is used for 55 estimates.

Estimated modelVals SWE-benchReact Native EvalsSource
Claude Fable 595.0%100.0%estimated ± 2.4 pp, low confidence
Claude Haiku 4.566.6%60.9%estimated ± 2.4 pp, low confidence
Claude Opus 4.888.6%97.0%estimated ± 2.4 pp, low confidence
Claude Opus 597.0%100.0%estimated ± 2.4 pp, low confidence
Claude Sonnet 579.6%80.6%estimated ± 2.4 pp, medium confidence
Composer 2.579.6%80.6%estimated ± 2.4 pp, medium confidence
DeepSeek V4 Flash 073188.8%97.4%estimated ± 2.4 pp, low confidence
DeepSeek V4 Pro 081396.4%100.0%estimated ± 2.4 pp, low confidence
Gemini 2.5 Pro54.4%45.6%estimated ± 2.4 pp, low confidence
Gemini 3.1 Flash-Lite62.8%55.8%estimated ± 2.4 pp, low confidence
Gemini 3.5 Flash78.8%79.3%estimated ± 2.4 pp, medium confidence
Gemini 3.5 Flash-Lite75.0%73.2%estimated ± 2.4 pp, medium confidence
Gemini 3.6 Flash79.6%80.6%estimated ± 2.4 pp, medium confidence
Gemini 3.7 Flash80.8%82.6%estimated ± 2.4 pp, medium confidence
Gemini 3.8 Flash80.0%81.3%estimated ± 2.4 pp, medium confidence
Gemini 3 Flash75.0%73.2%estimated ± 2.4 pp, medium confidence
GLM-4.769.4%64.8%estimated ± 2.4 pp, low confidence
GLM-5.176.4%75.4%estimated ± 2.4 pp, medium confidence
GLM-5.282.8%86.1%estimated ± 2.4 pp, low confidence
GLM-5.395.4%100.0%estimated ± 2.4 pp, low confidence
GLM-5.3-Flash92.0%100.0%estimated ± 2.4 pp, low confidence
GPT-5.2-Codex72.4%69.2%estimated ± 2.4 pp, low confidence
GPT-5.3 Codex78.0%78.0%estimated ± 2.4 pp, medium confidence
GPT-5.4 mini73.0%70.1%estimated ± 2.4 pp, low confidence
GPT-5.4 nano69.8%65.4%estimated ± 2.4 pp, low confidence
GPT-5.6 Luna93.0%100.0%estimated ± 2.4 pp, low confidence
GPT-5.6 Sol96.2%100.0%estimated ± 2.4 pp, low confidence
GPT-5.6 Terra95.4%100.0%estimated ± 2.4 pp, low confidence
Grok 4.2072.2%68.9%estimated ± 2.4 pp, low confidence
Grok 4.371.4%67.7%estimated ± 2.4 pp, low confidence
Grok 4.586.6%93.1%estimated ± 2.4 pp, low confidence
Grok 4.695.6%100.0%estimated ± 2.4 pp, low confidence
Inkling77.6%77.3%estimated ± 2.4 pp, medium confidence
Inkling-Small82.2%85.1%estimated ± 2.4 pp, medium confidence
Kimi K2.676.2%75.0%estimated ± 2.4 pp, medium confidence
Kimi K2.7 Code78.2%78.3%estimated ± 2.4 pp, medium confidence
Kimi K393.4%100.0%estimated ± 2.4 pp, low confidence
Laguna M.157.6%49.3%estimated ± 2.4 pp, low confidence
Laguna XS.255.2%46.5%estimated ± 2.4 pp, low confidence
Ling 3.0 Flash65.2%59.0%estimated ± 2.4 pp, low confidence
MiMo-V2.571.0%67.1%estimated ± 2.4 pp, low confidence
MiMo-V2.5-Pro74.0%71.6%estimated ± 2.4 pp, medium confidence
MiniMax M375.0%73.2%estimated ± 2.4 pp, medium confidence
Mistral Medium 3.5 128B66.4%60.6%estimated ± 2.4 pp, low confidence
Muse Spark74.4%72.2%estimated ± 2.4 pp, medium confidence
Muse Spark 1.182.0%84.7%estimated ± 2.4 pp, medium confidence
Muse Spark 1.286.6%93.1%estimated ± 2.4 pp, low confidence
Nemotron 3 Ultra69.0%64.2%estimated ± 2.4 pp, low confidence
Qwen3.5 Flash64.4%57.9%estimated ± 2.4 pp, low confidence
Qwen3.6-27B70.0%65.7%estimated ± 2.4 pp, low confidence
Qwen 3.6 Max (preview)72.8%69.8%estimated ± 2.4 pp, low confidence
Qwen3.6 Plus73.4%70.7%estimated ± 2.4 pp, low confidence
Qwen3.7 Max68.8%63.9%estimated ± 2.4 pp, low confidence
Qwen3.8-27B86.0%92.0%estimated ± 2.4 pp, low confidence
Qwen3.8 Max85.6%91.2%estimated ± 2.4 pp, low confidence