benchgap
Knowledge & reasoning

LABBench2 leaderboard

As of 2026-10-07, the highest measured score on LABBench2 is 88.8% by Gemini 4 Argon. 205 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 4 Argon88.8%measured
2Gemini 3.8 Flash86.2%measured
3Claude Sonnet 5.586.2%estimated ± 2.4 pp, medium confidence
4Claude Fable 585.6%estimated ± 1.3 pp, medium confidence
5GPT-6 Astra85.6%estimated ± 1.3 pp, medium confidence
6Claude Opus 584.2%measured
7Claude Fable 5.184.1%estimated ± 1.3 pp, medium confidence
8Claude Opus 5.584.1%estimated ± 1.3 pp, medium confidence
9MiMo-V2.6-Pro83.8%estimated ± 2.4 pp, medium confidence
10Gemini 3.1 Pro83.6%estimated ± 1.8 pp, low confidence
11Claude Opus 4.783.0%estimated ± 1.8 pp, low confidence
12Qwen3.7 Max82.7%estimated ± 1.8 pp, low confidence
13Muse Spark 1.382.7%estimated ± 1.8 pp, low confidence
14GPT-6.1 Sol82.6%estimated ± 1.3 pp, medium confidence
15Step 5 Preview82.6%estimated ± 2.4 pp, medium confidence
16Muse Spark 1.182.4%estimated ± 1.8 pp, low confidence
17Gemini 3 Flash82.3%estimated ± 1.8 pp, low confidence
18Qwen3.8 Max82.3%estimated ± 1.8 pp, low confidence
19Muse Spark 1.282.2%estimated ± 1.8 pp, low confidence
20Gemini 3.7 Flash82.1%measured
21GPT-5.6 Sol82.1%measured
22Qwen3.6 Plus81.8%estimated ± 1.8 pp, low confidence
23Kimi K2.681.8%estimated ± 1.8 pp, low confidence
24Claude Haiku 5.581.7%estimated ± 2.4 pp, medium confidence
25Muse Spark81.6%estimated ± 1.8 pp, low confidence
26Qwen3.8 Max Preview81.5%estimated ± 1.8 pp, low confidence
27GLM-5.181.4%estimated ± 1.8 pp, low confidence
28GLM-5.381.3%estimated ± 1.8 pp, low confidence
29GLM-5.281.3%estimated ± 1.8 pp, low confidence
30GPT-5.6 Terra81.2%measured
31Grok 4.781.1%estimated ± 2.4 pp, medium confidence
32Grok 4.2081.1%estimated ± 1.8 pp, low confidence
33Inkling81.1%estimated ± 1.8 pp, low confidence
34Gemini 3.1 Flash-Lite81.0%estimated ± 1.8 pp, low confidence
35GPT-6 Sol81.0%estimated ± 1.3 pp, medium confidence
36GLM-5.3-Flash80.9%estimated ± 1.8 pp, low confidence
37Gemini 3.5 Flash-Lite80.8%estimated ± 1.8 pp, low confidence
38Grok 4.380.8%estimated ± 1.8 pp, low confidence
39Nemotron 3 Ultra80.8%estimated ± 1.8 pp, low confidence
40Qwen3.8-Flash-Next80.7%estimated ± 1.8 pp, low confidence
41dots3-note Preview80.7%estimated ± 2.5 pp, low confidence
42Claude Sonnet 4.580.7%estimated ± 2.5 pp, low confidence
43Gemini 3 Pro Deep Think80.7%estimated ± 2.5 pp, low confidence
44GPT-5.580.2%estimated ± 1.3 pp, low confidence
45GPT-5.5 Pro80.2%estimated ± 1.3 pp, low confidence
46MiMo-V2.5-Pro80.1%estimated ± 1.8 pp, low confidence
47Claude Sonnet 580.1%measured
48Qwen3.8-27B79.9%estimated ± 1.8 pp, low confidence
49MiniMax M379.9%estimated ± 1.8 pp, low confidence
50Qwen3.5 Flash79.8%estimated ± 1.8 pp, low confidence
51Ling 3.1 Flash79.5%estimated ± 2.4 pp, low confidence
52DeepSeek V4.1 Flash79.5%estimated ± 2.4 pp, low confidence
53GPT-5.3 Codex79.4%estimated ± 1.8 pp, low confidence
54GPT-5.4 Pro79.3%estimated ± 1.3 pp, low confidence
55Kimi K379.3%estimated ± 1.3 pp, low confidence
56Claude Opus 4.7 (Adaptive)79.3%estimated ± 1.8 pp, low confidence
57MiMo-V2.579.1%estimated ± 1.8 pp, low confidence
58GLM-4.779.0%estimated ± 1.8 pp, low confidence
59GLM-4.678.7%estimated ± 1.8 pp, low confidence
60Ling 3.0 Flash78.6%estimated ± 1.8 pp, low confidence
61GLM-4.578.1%estimated ± 1.8 pp, low confidence
62GPT-5.478.0%estimated ± 1.3 pp, low confidence
63MiMo-V2.6-Flash77.7%estimated ± 2.4 pp, low confidence
64MiniMax M2.777.7%estimated ± 1.8 pp, low confidence
65Mistral Large 477.7%estimated ± 2.4 pp, low confidence
66Qwen3.7 Plus77.1%estimated ± 1.8 pp, low confidence
67GPT-5.2-Codex76.9%estimated ± 1.8 pp, low confidence
68Claude Opus 4.6 (Adaptive)76.8%estimated ± 1.3 pp, low confidence
69Claude Haiku 4.576.7%estimated ± 1.8 pp, low confidence
70Hy376.6%estimated ± 1.8 pp, low confidence
71Hy3 Preview76.6%estimated ± 1.8 pp, low confidence
72Kimi K2.7 Code76.4%estimated ± 1.8 pp, low confidence
73Claude Opus 4.875.9%estimated ± 1.3 pp, low confidence
74Gemini 3.5 Flash75.9%estimated ± 1.3 pp, low confidence
75Solar Pro 475.7%estimated ± 1.8 pp, low confidence
76Qwen 3.6 Max (preview)75.2%estimated ± 1.8 pp, low confidence
77Mistral Medium 3.5 128B74.6%estimated ± 1.8 pp, low confidence
78Kimi K2.573.8%estimated ± 1.8 pp, low confidence
79Kimi K2.5 (Reasoning)73.8%estimated ± 1.8 pp, low confidence
80Gemini 3.6 Flash73.6%estimated ± 1.3 pp, low confidence
81Grok 473.5%estimated ± 1.8 pp, low confidence
82MiMo-V2-Pro72.5%estimated ± 1.8 pp, low confidence
83Apodex 1.171.6%estimated ± 1.8 pp, low confidence
84Apodex 1.1 Mini71.6%estimated ± 1.8 pp, low confidence
85DeepSeek V4 Pro 081371.4%estimated ± 1.3 pp, low confidence
86Ling 3.0 Flash VL71.3%estimated ± 1.8 pp, low confidence
87Qwen3.5 397B71.2%estimated ± 1.8 pp, low confidence
88Qwen3.5 397B (Reasoning)71.2%estimated ± 1.8 pp, low confidence
89GPT-5.1-Codex71.0%estimated ± 1.8 pp, low confidence
90GPT-5.1-Codex-Max71.0%estimated ± 1.8 pp, low confidence
91Qwen3.5-27B70.8%estimated ± 1.8 pp, low confidence
92A.X K270.6%estimated ± 1.8 pp, low confidence
93Gemma 4 31B70.6%estimated ± 1.8 pp, low confidence
94Qwen3.5-122B-A10B70.6%estimated ± 1.8 pp, low confidence
95Laguna XS.270.6%estimated ± 1.8 pp, low confidence
96Laguna M.170.4%estimated ± 1.8 pp, low confidence
97Ling 3.0 Flash FP870.3%estimated ± 1.8 pp, low confidence
98GPT-5 (high)70.2%estimated ± 1.8 pp, low confidence
99Grok 4.1 Fast (Reasoning)70.0%estimated ± 1.8 pp, low confidence
100DeepSeek V4 Flash 073169.6%estimated ± 1.3 pp, low confidence
101GLM-5-Turbo69.2%estimated ± 1.8 pp, low confidence
102Grok 4 Fast (Reasoning)69.2%estimated ± 1.8 pp, low confidence
103o3-pro68.9%estimated ± 1.8 pp, low confidence
104Qwen3.5-35B-A3B68.9%estimated ± 1.8 pp, low confidence
105Gemini 2.5 Pro68.8%estimated ± 1.8 pp, low confidence
106GPT-5 (medium)68.5%estimated ± 1.8 pp, low confidence
107Qwen3.6-27B68.5%estimated ± 1.8 pp, low confidence
108Qwen3.6-35B-A3B68.3%estimated ± 1.8 pp, low confidence
109Claude Opus 4.668.2%estimated ± 1.8 pp, low confidence
110Mercury 2.567.8%estimated ± 2.4 pp, low confidence
111GPT-5.6 Luna67.6%estimated ± 1.3 pp, low confidence
112Muse Glimmer 30B67.5%estimated ± 1.8 pp, low confidence
113K-EXAONE 2.066.7%estimated ± 1.8 pp, low confidence
114MiMo-V2-Omni66.5%estimated ± 1.8 pp, low confidence
115o366.4%estimated ± 1.8 pp, low confidence
116Grok 4.665.7%estimated ± 1.3 pp, low confidence
117GLM-565.4%estimated ± 1.8 pp, low confidence
118GPT-6 Luna65.1%estimated ± 1.3 pp, low confidence
119DeepSeek-R164.5%estimated ± 1.8 pp, low confidence
120Claude Opus 4.564.1%estimated ± 1.8 pp, low confidence
121GPT-4 Turbo64.1%estimated ± 2.4 pp, low confidence
122GPT-5.264.1%estimated ± 1.3 pp, low confidence
123Claude 4.1 Opus Thinking64.0%estimated ± 1.8 pp, low confidence
124GLM-5V-Turbo64.0%estimated ± 1.8 pp, low confidence
125Step 3.7 Flash64.0%estimated ± 1.8 pp, low confidence
126Claude Sonnet 4.663.7%estimated ± 1.3 pp, low confidence
127Grok 4.563.1%estimated ± 1.3 pp, low confidence
128Nemotron 3 Super 100B62.8%estimated ± 1.8 pp, low confidence
129Gemma 4 26B A4B61.7%estimated ± 1.8 pp, low confidence
130K-Exaone60.6%estimated ± 1.8 pp, low confidence
131GPT-OSS 120B60.5%estimated ± 1.8 pp, low confidence
132DeepSeek V3.1 (Reasoning)60.1%estimated ± 1.8 pp, low confidence
133Inkling-Small59.7%estimated ± 1.3 pp, low confidence
134Mistral Small 458.8%estimated ± 1.8 pp, low confidence
135Mistral Small 4 (Reasoning)58.8%estimated ± 1.8 pp, low confidence
136Kimi K258.5%estimated ± 1.8 pp, low confidence
137o1-preview58.3%estimated ± 1.8 pp, low confidence
138Qwen3 Max58.2%estimated ± 1.8 pp, low confidence
139Command A+57.8%estimated ± 1.8 pp, low confidence
140Nemotron 3 Nano 30B57.4%estimated ± 1.8 pp, low confidence
141North Mini Code57.4%estimated ± 1.8 pp, low confidence
142Gemma 4 12B56.9%estimated ± 1.8 pp, low confidence
143Trinity-Large-Preview56.7%estimated ± 1.8 pp, low confidence
144Trinity-Large-Thinking56.7%estimated ± 1.8 pp, low confidence
145DeepSeek V3.256.6%estimated ± 1.8 pp, low confidence
146o3-mini56.3%estimated ± 1.8 pp, low confidence
147o156.1%estimated ± 1.8 pp, low confidence
148Nemotron 3.5 Lightning 30B A3B NVFP455.7%estimated ± 1.8 pp, low confidence
149Sarvam 105B55.1%estimated ± 1.8 pp, low confidence
150DeepSeek V3.154.7%estimated ± 1.8 pp, low confidence
151Ling 3.0 Tiny54.6%estimated ± 1.8 pp, low confidence
152GLM-4.5-Air54.5%estimated ± 1.8 pp, low confidence
153Quasar 438B54.4%estimated ± 1.8 pp, low confidence
154Nemotron Ultra 253B53.9%estimated ± 1.8 pp, low confidence
155Grok Code Fast 153.8%estimated ± 1.8 pp, low confidence
156Qwen3-Omni-30B-A3B-Thinking53.7%estimated ± 1.8 pp, low confidence
157Solar Pro 353.4%estimated ± 1.8 pp, low confidence
158Claude Opus 4.5 Thinking51.6%estimated ± 1.3 pp, low confidence
159MiniCPM5-2B50.9%estimated ± 1.8 pp, low confidence
160GPT-OSS 20B49.4%estimated ± 1.8 pp, low confidence
161Claude 4 Sonnet48.8%estimated ± 1.8 pp, low confidence
162Gemini 2.5 Flash48.8%estimated ± 1.8 pp, low confidence
163Mistral Large 348.5%estimated ± 1.8 pp, low confidence
164Llama 4 Maverick47.5%estimated ± 1.8 pp, low confidence
165GPT-4.147.0%estimated ± 1.8 pp, low confidence
166GPT-4.1 mini46.8%estimated ± 1.8 pp, low confidence
167MiMo-V2-Flash46.0%estimated ± 1.8 pp, low confidence
168DeepSeek V3 032445.9%estimated ± 1.8 pp, low confidence
169Granite 4.2 30B44.7%estimated ± 1.8 pp, low confidence
170Grok 4.1 Fast44.0%estimated ± 1.8 pp, low confidence
171Sarvam 30B43.6%estimated ± 1.8 pp, low confidence
172Celeris-143.4%estimated ± 1.8 pp, low confidence
173Granite 4.2 8B43.4%estimated ± 1.8 pp, low confidence
174Exaone 4.0 32B43.1%estimated ± 1.8 pp, low confidence
175Qwen3-Omni-30B-A3B-Instruct42.3%estimated ± 1.8 pp, low confidence
176DeepSeek R1 Distill Qwen 32B41.8%estimated ± 1.8 pp, low confidence
177Gemini 3 Pro41.6%estimated ± 1.3 pp, low confidence
178Ling 2.6 Flash39.7%estimated ± 1.8 pp, low confidence
179Gemini 1.5 Pro39.3%estimated ± 1.8 pp, low confidence
180Llama 4 Scout39.1%estimated ± 1.8 pp, low confidence
181Mistral Medium 338.3%estimated ± 1.8 pp, low confidence
182Gemma 4 E4B38.1%estimated ± 1.8 pp, low confidence
183Phi-438.0%estimated ± 1.8 pp, low confidence
184GPT-5.137.4%estimated ± 1.3 pp, low confidence
185Solar Pro 236.7%estimated ± 1.8 pp, low confidence
186Granite 4.2 3B36.5%estimated ± 1.8 pp, low confidence
187LFM2.5-2.6B36.4%estimated ± 1.8 pp, low confidence
188DeepSeek V336.4%estimated ± 1.8 pp, low confidence
189GPT-4o35.1%estimated ± 1.8 pp, low confidence
190Llama 3.1 405B32.7%estimated ± 1.8 pp, low confidence
191LFM2.5-8B-A1B32.5%estimated ± 1.8 pp, low confidence
192GPT-4.1 nano32.4%estimated ± 1.8 pp, low confidence
193Nova Pro31.3%estimated ± 1.8 pp, low confidence
194Ultravox v0.6 Llama 3.3 70B31.2%estimated ± 1.8 pp, low confidence
195Claude 3 Opus30.5%estimated ± 1.8 pp, low confidence
196Mistral Large 230.2%estimated ± 1.8 pp, low confidence
197Nemotron 3 Nano Omni 30B A3B28.9%estimated ± 1.8 pp, low confidence
198Gemma 4 E2B26.0%estimated ± 1.8 pp, low confidence
199Gemma 3 27B25.6%estimated ± 1.8 pp, low confidence
200GPT-4o mini25.5%estimated ± 1.8 pp, low confidence
201Exaone 4.0 1.2B25.3%estimated ± 1.8 pp, low confidence
202Qwen2.5 Coder 32B Instruct24.8%estimated ± 1.8 pp, low confidence
203GPT-5.4 mini22.8%estimated ± 1.3 pp, low confidence
204Claude Sonnet 4.5 Thinking22.7%estimated ± 1.3 pp, low confidence
205Claude 3 Haiku21.7%estimated ± 1.8 pp, low confidence
206Phi-4 Multimodal Instruct17.6%estimated ± 1.8 pp, low confidence
207LFM2.5-VL-1.6B-Extract15.9%estimated ± 1.8 pp, low confidence
208Gemini 1.0 Pro15.1%estimated ± 1.8 pp, low confidence
209Granite-4.0-H-1B14.3%estimated ± 1.8 pp, low confidence
210Granite-4.0-350M14.1%estimated ± 1.8 pp, low confidence
211Granite-4.0-H-350M13.9%estimated ± 1.8 pp, low confidence
212GPT-5.4 nano11.5%estimated ± 1.3 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General