benchgap
Coding

NL2Repo leaderboard

As of 2026-10-07, the highest measured score on NL2Repo is 65.4% by DeepSeek V4.1 Flash. 168 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.581.4%estimated ± 6.4 pp, low confidence
2Claude Mythos 576.7%estimated ± 6.1 pp, low confidence
3Claude Sonnet 5.571.6%estimated ± 6.4 pp, low confidence
4Ember-171.2%estimated ± 6.1 pp, low confidence
5Gemini 4 Argon68.6%estimated ± 6.5 pp, low confidence
6MiMo-V2.6-Pro67.2%estimated ± 6.5 pp, low confidence
7DeepSeek V4.1 Flash65.4%measured
8Step 5 Preview64.0%estimated ± 6.5 pp, medium confidence
9GPT-6 Sol61.9%estimated ± 6.5 pp, medium confidence
10Grok 4.761.6%estimated ± 6.5 pp, medium confidence
11DeepSeek V4 Pro 081361.5%measured
12Pareto 26.961.3%estimated ± 6.6 pp, medium confidence
13Claude Opus 560.9%estimated ± 4.7 pp, medium confidence
14SWE-260.9%estimated ± 6.6 pp, medium confidence
15Claude Fable 5.160.7%estimated ± 5.1 pp, low confidence
16GPT-5.6 Sol60.1%estimated ± 4.7 pp, high confidence
17Pareto 26.10 Preview59.7%estimated ± 6.6 pp, medium confidence
18Ornith-1.5-397B59.5%measured
19Grok 4.659.4%estimated ± 4.7 pp, high confidence
20GPT-5.6 Terra59.2%estimated ± 4.7 pp, high confidence
21Hy4 preview58.9%measured
22Claude Fable 558.8%estimated ± 4.7 pp, high confidence
23GLM-5.358.0%measured
24Claude Haiku 5.557.8%estimated ± 6.5 pp, medium confidence
25Kimi K357.1%estimated ± 4.7 pp, high confidence
26GPT-6 Luna57.1%estimated ± 6.5 pp, medium confidence
27GPT-6 Astra57.1%estimated ± 5.1 pp, low confidence
28GPT-5.6 Luna56.7%estimated ± 4.7 pp, high confidence
29GPT-6.1 Sol56.5%estimated ± 6.5 pp, medium confidence
30Mistral Large 456.5%estimated ± 6.5 pp, medium confidence
31Ling 3.1 Flash56.3%estimated ± 6.5 pp, medium confidence
32GLM-5.3-Flash56.3%measured
33Muse Spark 1.356.2%estimated ± 5.1 pp, low confidence
34Qwen3.8 Max55.9%measured
35Claude Opus 4.7 (Adaptive)54.5%estimated ± 5.1 pp, medium confidence
36DeepSeek V4 Flash 073154.2%measured
37Qwen3.8 Max Preview53.1%estimated ± 5.1 pp, medium confidence
38Claude Opus 4.852.5%estimated ± 4.7 pp, high confidence
39GPT-5.452.5%estimated ± 5.1 pp, medium confidence
40MiMo-V2.6-Flash51.8%estimated ± 6.5 pp, medium confidence
41Grok 4.550.8%estimated ± 4.7 pp, high confidence
42Muse Spark 1.250.8%estimated ± 4.7 pp, high confidence
43dots3-note Preview49.8%measured
44Beam49.2%estimated ± 6.1 pp, medium confidence
45Claude Opus 4.649.0%estimated ± 6.1 pp, medium confidence
46GLM-5.248.9%measured
47Qwen3.8-Omni-Flash48.9%measured
48Ornith-1.0-397B48.2%measured
49Qwen3.8-Flash-Next48.1%measured
50Sakana Fugu-Ultra47.7%estimated ± 3.1 pp, low confidence
51GPT-5.547.7%estimated ± 4.7 pp, high confidence
52Sakana Fugu47.5%estimated ± 3.1 pp, low confidence
53Inkling-Small47.4%estimated ± 4.7 pp, high confidence
54Solar Open 247.3%estimated ± 3.1 pp, low confidence
55GPT-5.247.3%estimated ± 6.1 pp, medium confidence
56Claude Opus 4.747.2%estimated ± 4.7 pp, high confidence
57Muse Spark 1.147.2%estimated ± 4.7 pp, high confidence
58Qwen3.7 Max47.2%measured
59Atria Dawn Preview47.0%estimated ± 6.4 pp, medium confidence
60Seed 2.1 Pro47.0%measured
61Laguna S 2.146.8%estimated ± 6.4 pp, medium confidence
62Ternary Bonsai 2 27B46.4%estimated ± 3.1 pp, low confidence
63Gemini 3.7 Flash46.4%estimated ± 4.7 pp, high confidence
64Kimi K2.646.2%estimated ± 3.1 pp, low confidence
65Ornith-1.5-35B-A3B46.2%measured
66Nemotron 3 Ultra46.0%estimated ± 3.1 pp, low confidence
67Gemini 3.8 Flash45.9%estimated ± 4.7 pp, high confidence
68BTL-345.6%estimated ± 3.1 pp, low confidence
69Claude Sonnet 545.6%estimated ± 4.7 pp, high confidence
70Composer 2.545.6%estimated ± 4.7 pp, high confidence
71Gemini 3.6 Flash45.6%estimated ± 4.7 pp, high confidence
72MAI-Thinking-145.5%estimated ± 3.1 pp, low confidence
73Qwen3.6 Plus45.2%estimated ± 3.1 pp, low confidence
74Gemini 3.1 Pro45.1%estimated ± 4.7 pp, high confidence
75Gemini 3.5 Flash45.1%estimated ± 4.7 pp, high confidence
76Quasar 438B44.8%estimated ± 5.1 pp, medium confidence
77Kimi K2.7 Code44.7%estimated ± 4.7 pp, high confidence
78GPT-5.3 Codex44.6%estimated ± 4.7 pp, high confidence
79Apodex 1.144.5%estimated ± 5.1 pp, medium confidence
80Apodex 1.1 Mini44.5%estimated ± 5.1 pp, medium confidence
81Kimi K2.544.4%estimated ± 3.1 pp, low confidence
82Inkling44.4%estimated ± 4.7 pp, high confidence
83Claude Sonnet 4.644.3%estimated ± 4.7 pp, high confidence
84A.X K244.0%estimated ± 3.1 pp, low confidence
85Qwen3.5 397B43.8%estimated ± 3.1 pp, low confidence
86Seed 2.1 Turbo43.7%measured
87MiMo-V2-Pro43.2%estimated ± 6.1 pp, medium confidence
88Claude Opus 4.543.2%measured
89Gemini 3.5 Flash-Lite43.0%estimated ± 4.7 pp, high confidence
90Gemini 3 Flash43.0%estimated ± 4.7 pp, high confidence
91Hy343.0%estimated ± 5.1 pp, medium confidence
92Hy3 Preview43.0%estimated ± 5.1 pp, medium confidence
93Qwen 3.6 Max (preview)42.9%measured
94GLM-542.8%estimated ± 6.1 pp, medium confidence
95GLM-5.142.7%measured
96Muse Spark42.7%estimated ± 4.7 pp, high confidence
97MiMo-V2.5-Pro42.5%estimated ± 4.7 pp, high confidence
98Qwen3.8-27B42.3%measured
99MiniMax M342.1%measured
100GPT-5.4 mini42.0%estimated ± 4.7 pp, high confidence
101GPT-5.2-Codex41.7%estimated ± 4.7 pp, high confidence
102Grok 4.2041.6%estimated ± 4.7 pp, high confidence
103Claude Sonnet 4.541.6%estimated ± 6.1 pp, medium confidence
104Grok 4.341.3%estimated ± 4.7 pp, high confidence
105MiMo-V2.541.1%estimated ± 4.7 pp, high confidence
106Qwen3.7 Plus41.1%measured
107GPT-5.4 nano40.7%estimated ± 4.7 pp, high confidence
108Granite 4.2 30B40.5%estimated ± 3.1 pp, low confidence
109GLM-4.740.5%estimated ± 4.7 pp, high confidence
110Ling 3.0 Flash VL40.5%estimated ± 6.5 pp, medium confidence
111MiniMax M2.739.8%measured
112Claude Haiku 4.539.6%estimated ± 4.7 pp, medium confidence
113Mistral Medium 3.5 128B39.5%estimated ± 4.7 pp, medium confidence
114Granite 4.2 8B39.5%estimated ± 3.1 pp, low confidence
115Ling 3.0 Flash39.2%estimated ± 4.7 pp, medium confidence
116Qwen3.5 Flash38.9%estimated ± 4.7 pp, medium confidence
117Gemma 4 12B38.9%estimated ± 3.1 pp, low confidence
118Gemini 3.1 Flash-Lite38.5%estimated ± 4.7 pp, medium confidence
119Mellum2-12B-A2.5B-Thinking38.0%estimated ± 3.1 pp, low confidence
120Granite 4.2 3B37.9%estimated ± 3.1 pp, low confidence
121MiniCPM5-2B37.7%estimated ± 3.1 pp, low confidence
122Laguna M.137.5%estimated ± 4.7 pp, medium confidence
123Laguna XS.237.2%estimated ± 4.7 pp, medium confidence
124Gemini 2.5 Pro37.1%estimated ± 4.7 pp, medium confidence
125MiMo-V2-Omni36.8%estimated ± 6.1 pp, medium confidence
126Ling 3.0 Flash FP836.6%estimated ± 5.1 pp, medium confidence
127BTL-436.3%estimated ± 3.1 pp, low confidence
128Claude 4.1 Opus36.2%estimated ± 6.1 pp, medium confidence
129Qwen3.6-27B36.2%measured
130ZAYA1-8B36.2%estimated ± 3.1 pp, low confidence
131ZAYA1-74B-Preview36.1%estimated ± 3.1 pp, low confidence
132MiMo-V2-Flash36.0%estimated ± 5.1 pp, medium confidence
133GPT-5.135.6%estimated ± 5.1 pp, medium confidence
134Muse Glimmer 30B35.3%estimated ± 5.1 pp, medium confidence
135Ornith-1.0-35B34.6%measured
136Kimi K2.5 (Reasoning)33.6%estimated ± 5.1 pp, medium confidence
137Agents-A1-4B33.3%estimated ± 3.1 pp, low confidence
138LFM2.5-2.6B33.2%estimated ± 3.1 pp, low confidence
139Claude 4 Sonnet32.8%estimated ± 6.1 pp, medium confidence
140Qwen3.5-122B-A10B32.8%estimated ± 5.1 pp, medium confidence
141MAI-Code-1.1-Flash32.6%estimated ± 6.1 pp, medium confidence
142Ornith-1.5-9B32.4%measured
143Qwen3.5-27B32.2%estimated ± 6.1 pp, medium confidence
144DeepSeek V3 032432.2%estimated ± 6.5 pp, medium confidence
145North Mini Code31.8%estimated ± 6.5 pp, medium confidence
146Mercury 2.531.4%estimated ± 6.5 pp, medium confidence
147Gemma 4 31B31.0%estimated ± 5.1 pp, medium confidence
148Laguna XS 2.129.5%estimated ± 6.1 pp, medium confidence
149Qwen3.6-35B-A3B29.4%measured
150Grok Code Fast 129.3%estimated ± 6.1 pp, medium confidence
151Solar Pro 428.9%estimated ± 6.1 pp, medium confidence
152o128.1%estimated ± 5.1 pp, low confidence
153Step 3.7 Flash28.0%estimated ± 5.1 pp, low confidence
154Gemma 4 26B A4B27.8%estimated ± 5.1 pp, low confidence
155Ornith-1.0-9B27.2%measured
156GPT-5 (high)26.6%estimated ± 5.1 pp, low confidence
157Nemotron 3 Super 100B26.5%estimated ± 5.1 pp, low confidence
158Qwen3.5-35B-A3B26.5%estimated ± 6.1 pp, low confidence
159K-EXAONE 2.024.8%estimated ± 6.1 pp, low confidence
160LongCat-Flash-Lite-Sparse24.8%estimated ± 6.1 pp, low confidence
161o1-preview23.7%estimated ± 5.1 pp, low confidence
162Mellum2-12B-A2.5B-Instruct22.2%estimated ± 3.1 pp, low confidence
163K-Exaone22.2%estimated ± 5.1 pp, low confidence
164GPT-OSS 120B20.9%estimated ± 5.1 pp, low confidence
165MiniCPM5-1B20.2%estimated ± 3.1 pp, low confidence
166Command A+18.8%estimated ± 5.1 pp, low confidence
167Nemotron 3.5 Lightning 30B A3B NVFP418.0%estimated ± 5.1 pp, low confidence
168Mistral Small 417.9%estimated ± 5.1 pp, low confidence
169Mistral Small 4 (Reasoning)17.9%estimated ± 5.1 pp, low confidence
170LLaDA2.2-mini17.3%estimated ± 3.1 pp, low confidence
171Trinity-Large-Preview17.2%estimated ± 5.1 pp, low confidence
172Trinity-Large-Thinking17.2%estimated ± 5.1 pp, low confidence
173Ling 2.6 Flash16.8%estimated ± 5.1 pp, low confidence
174Gemini 1.5 Pro15.6%estimated ± 5.1 pp, low confidence
175DeepSeek V315.1%estimated ± 5.1 pp, low confidence
176GPT-4 Turbo13.9%estimated ± 5.1 pp, low confidence
177GPT-OSS 20B13.3%estimated ± 5.1 pp, low confidence
178GPT-4.1 mini12.9%estimated ± 5.1 pp, low confidence
179Mistral Large 312.8%estimated ± 5.1 pp, low confidence
180Claude 3 Opus12.4%estimated ± 5.1 pp, low confidence
181Solar Pro 310.6%estimated ± 6.5 pp, low confidence
182Llama 4 Maverick9.8%estimated ± 5.1 pp, low confidence
183GPT-4.18.6%estimated ± 6.1 pp, low confidence
184Ling 3.0 Tiny8.5%estimated ± 6.5 pp, low confidence
185Celeris-18.4%estimated ± 5.1 pp, low confidence
186Nemotron 3 Nano 30B8.3%estimated ± 5.1 pp, low confidence
187Nemotron 3 Nano Omni 30B A3B7.9%estimated ± 5.1 pp, low confidence
188Ultravox v0.6 Llama 3.3 70B6.4%estimated ± 5.1 pp, low confidence
189GPT-4o mini6.0%estimated ± 5.1 pp, low confidence
190GPT-4.1 nano5.8%estimated ± 5.1 pp, low confidence
191o3-mini5.4%estimated ± 6.1 pp, low confidence
192LLaDA2.2-flash5.4%estimated ± 6.1 pp, low confidence
193Claude 3.5 Sonnet5.2%estimated ± 6.1 pp, low confidence
194Gemma 3 27B5.0%estimated ± 5.1 pp, low confidence
195Gemma 4 E4B4.5%estimated ± 5.1 pp, low confidence
196Llama 4 Scout3.5%estimated ± 5.1 pp, low confidence
197Gemma 4 E2B2.8%estimated ± 5.1 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General