benchgap
Coding

React Native Evals leaderboard

As of 2026-10-07, the highest measured score on React Native Evals is 96.1% by Composer 2. 151 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Fable 5100.0%estimated ± 2.4 pp, low confidence
2Claude Opus 5100.0%estimated ± 2.4 pp, low confidence
3DeepSeek V4 Pro 0813100.0%estimated ± 2.4 pp, low confidence
4GLM-5.3100.0%estimated ± 2.4 pp, low confidence
5GLM-5.3-Flash100.0%estimated ± 2.4 pp, low confidence
6GPT-5.6 Luna100.0%estimated ± 2.4 pp, low confidence
7GPT-5.6 Sol100.0%estimated ± 2.4 pp, low confidence
8GPT-5.6 Terra100.0%estimated ± 2.4 pp, low confidence
9Grok 4.6100.0%estimated ± 2.4 pp, low confidence
10Kimi K3100.0%estimated ± 2.4 pp, low confidence
11DeepSeek V4 Flash 073197.4%estimated ± 2.4 pp, low confidence
12Claude Opus 4.897.0%estimated ± 2.4 pp, low confidence
13Composer 296.1%measured
14Composer 2 Fast94.9%measured
15Grok 4.593.1%estimated ± 2.4 pp, low confidence
16Muse Spark 1.293.1%estimated ± 2.4 pp, low confidence
17Qwen3.8-27B92.0%estimated ± 2.4 pp, low confidence
18Qwen3.8 Max91.2%estimated ± 2.4 pp, low confidence
19Gemini 4 Argon86.5%estimated ± 3.0 pp, low confidence
20GLM-5.286.1%estimated ± 2.4 pp, low confidence
21Mistral Large 485.4%estimated ± 3.0 pp, low confidence
22GPT-5.485.3%measured
23Inkling-Small85.1%estimated ± 2.4 pp, medium confidence
24Muse Spark 1.184.7%estimated ± 2.4 pp, medium confidence
25GPT-5.584.7%measured
26Claude Fable 5.184.6%estimated ± 3.5 pp, medium confidence
27Claude Opus 4.684.1%measured
28Claude Opus 5.583.8%estimated ± 4.7 pp, low confidence
29GPT-6 Astra83.4%estimated ± 3.5 pp, medium confidence
30Muse Spark 1.383.1%estimated ± 3.5 pp, medium confidence
31Claude Opus 4.782.8%measured
32Gemini 3.7 Flash82.6%estimated ± 2.4 pp, medium confidence
33Claude Opus 4.7 (Adaptive)82.6%estimated ± 3.5 pp, high confidence
34Qwen3.8-Flash-Next82.4%estimated ± 3.5 pp, high confidence
35GPT-5.282.2%estimated ± 3.0 pp, medium confidence
36Claude Opus 4.6 (Adaptive)82.2%estimated ± 3.0 pp, medium confidence
37Qwen3.8 Max Preview82.1%estimated ± 3.5 pp, high confidence
38Claude Sonnet 5.581.3%estimated ± 4.7 pp, low confidence
39MiMo-V2.6-Pro81.3%estimated ± 4.7 pp, low confidence
40Gemini 3.8 Flash81.3%estimated ± 2.4 pp, medium confidence
41Claude Sonnet 580.6%estimated ± 2.4 pp, medium confidence
42Composer 2.580.6%estimated ± 2.4 pp, medium confidence
43Gemini 3.6 Flash80.6%estimated ± 2.4 pp, medium confidence
44Claude Sonnet 4.680.6%measured
45Step 5 Preview80.5%estimated ± 4.7 pp, low confidence
46GPT-6 Sol79.9%estimated ± 4.7 pp, medium confidence
47Grok 4.779.8%estimated ± 4.7 pp, medium confidence
48Quasar 438B79.5%estimated ± 3.5 pp, high confidence
49Apodex 1.179.3%estimated ± 3.5 pp, high confidence
50Apodex 1.1 Mini79.3%estimated ± 3.5 pp, high confidence
51Gemini 3.5 Flash79.3%estimated ± 2.4 pp, medium confidence
52Gemini 3.1 Pro78.9%measured
53Hy378.8%estimated ± 3.5 pp, high confidence
54Hy3 Preview78.8%estimated ± 3.5 pp, high confidence
55Claude Haiku 5.578.8%estimated ± 4.7 pp, medium confidence
56GPT-6 Luna78.6%estimated ± 4.7 pp, medium confidence
57GPT-6.1 Sol78.5%estimated ± 4.7 pp, medium confidence
58Ling 3.1 Flash78.4%estimated ± 4.7 pp, medium confidence
59Kimi K2.7 Code78.3%estimated ± 2.4 pp, medium confidence
60Qwen3.7 Plus78.1%estimated ± 3.5 pp, high confidence
61GPT-5.3 Codex78.0%estimated ± 2.4 pp, medium confidence
62DeepSeek V4.1 Flash77.5%estimated ± 4.7 pp, medium confidence
63Inkling77.3%estimated ± 2.4 pp, medium confidence
64MiMo-V2.6-Flash77.2%estimated ± 4.7 pp, medium confidence
65Kimi K2.577.2%measured
66Ling 3.0 Flash FP876.8%estimated ± 3.5 pp, high confidence
67MiMo-V2-Flash76.6%estimated ± 3.5 pp, high confidence
68Muse Glimmer 30B76.4%estimated ± 3.5 pp, high confidence
69Qwen3.5-122B-A10B75.6%estimated ± 3.5 pp, high confidence
70GLM-5.175.4%estimated ± 2.4 pp, medium confidence
71Gemma 4 31B75.2%measured
72Kimi K2.675.0%estimated ± 2.4 pp, medium confidence
73GLM-574.8%measured
74Qwen3.6-35B-A3B74.6%estimated ± 3.5 pp, high confidence
75Solar Pro 474.4%estimated ± 4.7 pp, medium confidence
76Ling 3.0 Flash VL74.2%estimated ± 4.7 pp, medium confidence
77o174.1%estimated ± 3.5 pp, high confidence
78Step 3.7 Flash74.0%estimated ± 3.5 pp, high confidence
79Gemma 4 26B A4B74.0%estimated ± 3.5 pp, high confidence
80Nemotron 3 Super 100B73.6%estimated ± 3.5 pp, high confidence
81K-EXAONE 2.073.3%estimated ± 4.7 pp, medium confidence
82Gemini 3.5 Flash-Lite73.2%estimated ± 2.4 pp, medium confidence
83Gemini 3 Flash73.2%estimated ± 2.4 pp, medium confidence
84MiniMax M373.2%estimated ± 2.4 pp, medium confidence
85A.X K272.8%estimated ± 4.7 pp, medium confidence
86o1-preview72.6%estimated ± 3.5 pp, high confidence
87Grok 472.6%measured
88Muse Spark72.2%estimated ± 2.4 pp, medium confidence
89K-Exaone72.1%estimated ± 3.5 pp, high confidence
90GPT-5.172.1%estimated ± 3.0 pp, low confidence
91DeepSeek V3 032472.0%estimated ± 4.7 pp, medium confidence
92North Mini Code71.9%estimated ± 4.7 pp, medium confidence
93Gemma 4 12B71.9%estimated ± 3.5 pp, high confidence
94Mercury 2.571.8%estimated ± 4.7 pp, medium confidence
95MiMo-V2.5-Pro71.6%estimated ± 2.4 pp, medium confidence
96GPT-OSS 120B71.6%measured
97DeepSeek V3.271.5%measured
98Granite 4.2 30B71.5%estimated ± 4.7 pp, medium confidence
99MiniMax M2.771.4%measured
100GLM-5 (Reasoning)71.2%estimated ± 3.0 pp, low confidence
101Command A+71.1%estimated ± 3.5 pp, high confidence
102GPT-OSS 20B71.0%measured
103Nemotron 3.5 Lightning 30B A3B NVFP470.8%estimated ± 3.5 pp, high confidence
104Mistral Small 470.8%estimated ± 3.5 pp, high confidence
105Mistral Small 4 (Reasoning)70.8%estimated ± 3.5 pp, high confidence
106Qwen3.6 Plus70.7%estimated ± 2.4 pp, low confidence
107Claude Sonnet 4.5 Thinking70.7%estimated ± 3.0 pp, low confidence
108Trinity-Large-Preview70.5%estimated ± 3.5 pp, high confidence
109Trinity-Large-Thinking70.5%estimated ± 3.5 pp, high confidence
110Ling 2.6 Flash70.4%estimated ± 3.5 pp, high confidence
111GPT-5.1-Codex-Max70.3%estimated ± 3.0 pp, low confidence
112GPT-5.4 mini70.1%estimated ± 2.4 pp, low confidence
113Gemini 1.5 Pro70.0%estimated ± 3.5 pp, high confidence
114DeepSeek V369.9%estimated ± 3.5 pp, high confidence
115Qwen 3.6 Max (preview)69.8%estimated ± 2.4 pp, low confidence
116Granite 4.2 8B69.7%estimated ± 3.5 pp, high confidence
117GPT-4 Turbo69.5%estimated ± 3.5 pp, high confidence
118GPT-5.2-Codex69.2%estimated ± 2.4 pp, low confidence
119GPT-4.1 mini69.2%estimated ± 3.5 pp, medium confidence
120Mistral Large 369.1%estimated ± 3.5 pp, medium confidence
121Claude Opus 4.5 Thinking69.0%estimated ± 3.0 pp, low confidence
122Claude 3 Opus69.0%estimated ± 3.5 pp, medium confidence
123Grok 4.2068.9%estimated ± 2.4 pp, low confidence
124GPT-5 (high)68.6%estimated ± 3.0 pp, low confidence
125Llama 4 Maverick68.2%estimated ± 3.5 pp, medium confidence
126Grok 4.367.7%estimated ± 2.4 pp, low confidence
127Celeris-167.7%estimated ± 3.5 pp, medium confidence
128Nemotron 3 Nano 30B67.7%estimated ± 3.5 pp, medium confidence
129Nemotron 3 Nano Omni 30B A3B67.5%estimated ± 3.5 pp, medium confidence
130MiMo-V2.567.1%estimated ± 2.4 pp, low confidence
131Ultravox v0.6 Llama 3.3 70B67.1%estimated ± 3.5 pp, medium confidence
132GPT-4o mini66.9%estimated ± 3.5 pp, medium confidence
133GPT-4.1 nano66.9%estimated ± 3.5 pp, medium confidence
134MiniCPM5-2B66.6%estimated ± 4.7 pp, low confidence
135Gemma 3 27B66.6%estimated ± 3.5 pp, medium confidence
136Gemma 4 E4B66.4%estimated ± 3.5 pp, medium confidence
137Solar Pro 366.3%estimated ± 4.7 pp, low confidence
138Granite 4.2 3B66.2%estimated ± 4.7 pp, low confidence
139Llama 4 Scout66.1%estimated ± 3.5 pp, medium confidence
140LFM2.5-2.6B66.0%estimated ± 3.5 pp, medium confidence
141Kimi K2.5 (Reasoning)66.0%estimated ± 3.0 pp, low confidence
142Gemma 4 E2B65.9%estimated ± 3.5 pp, medium confidence
143Ling 3.0 Tiny65.7%estimated ± 4.7 pp, low confidence
144Qwen3.6-27B65.7%estimated ± 2.4 pp, low confidence
145GPT-5.4 nano65.4%estimated ± 2.4 pp, low confidence
146GLM-4.764.8%estimated ± 2.4 pp, low confidence
147Nemotron 3 Ultra64.2%estimated ± 2.4 pp, low confidence
148Qwen3.7 Max63.9%estimated ± 2.4 pp, low confidence
149Qwen3.5 Plus63.9%estimated ± 3.0 pp, low confidence
150MiniMax M2.562.7%estimated ± 3.0 pp, low confidence
151Gemini 3 Pro61.9%estimated ± 3.0 pp, low confidence
152GPT-5 mini61.7%estimated ± 3.0 pp, low confidence
153Claude Haiku 4.560.9%estimated ± 2.4 pp, low confidence
154Mistral Medium 3.5 128B60.6%estimated ± 2.4 pp, low confidence
155GPT-5.1-Codex60.1%estimated ± 3.0 pp, low confidence
156Ling 3.0 Flash59.0%estimated ± 2.4 pp, low confidence
157Qwen3.5 Flash57.9%estimated ± 2.4 pp, low confidence
158Claude Haiku 4.5 Thinking57.0%estimated ± 3.0 pp, low confidence
159Gemini 3.1 Flash-Lite55.8%estimated ± 2.4 pp, low confidence
160Laguna M.149.3%estimated ± 2.4 pp, low confidence
161Laguna XS.246.5%estimated ± 2.4 pp, low confidence
162Gemini 2.5 Pro45.6%estimated ± 2.4 pp, low confidence
163DeepSeek V3.2 (Thinking)38.5%estimated ± 3.0 pp, low confidence
164Qwen3 Max30.4%estimated ± 3.0 pp, low confidence
165GLM-4.627.9%estimated ± 3.0 pp, low confidence
166Grok 4.1 Fast (Reasoning)13.2%estimated ± 3.0 pp, low confidence
167Grok 4 Fast (Reasoning)0.0%estimated ± 3.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General