benchgap
Long context

LongBench v2 leaderboard

As of 2026-10-07, the highest measured score on LongBench v2 is 66.3% by Qwen3.8 Max. 176 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.8 Max66.3%measured
2Beam65.5%measured
3Claude Opus 4.564.4%measured
4Qwen3.5 397B63.2%measured
5Qwen3.6 Plus62.0%measured
6Nemotron 3 Ultra61.9%measured
7A.X K261.5%estimated ± 5.9 pp, medium confidence
8Apodex 1.161.5%estimated ± 5.9 pp, low confidence
9Apodex 1.1 Mini61.5%estimated ± 5.9 pp, low confidence
10Claude 4.1 Opus Thinking61.5%estimated ± 5.9 pp, medium confidence
11Claude Fable 561.5%estimated ± 5.9 pp, low confidence
12Claude Fable 5.161.5%estimated ± 5.9 pp, low confidence
13Claude Haiku 5.561.5%estimated ± 5.9 pp, low confidence
14Claude Opus 4.5 Thinking61.5%estimated ± 5.9 pp, medium confidence
15Claude Opus 4.661.5%estimated ± 5.9 pp, medium confidence
16Claude Opus 4.6 (Adaptive)61.5%estimated ± 5.9 pp, medium confidence
17Claude Opus 4.761.5%estimated ± 5.9 pp, medium confidence
18Claude Opus 4.7 (Adaptive)61.5%estimated ± 5.9 pp, low confidence
19Claude Opus 4.861.5%estimated ± 5.9 pp, medium confidence
20Claude Opus 561.5%estimated ± 5.9 pp, low confidence
21Claude Opus 5.561.5%estimated ± 5.9 pp, low confidence
22Claude Sonnet 4.661.5%estimated ± 5.9 pp, medium confidence
23Claude Sonnet 561.5%estimated ± 5.9 pp, low confidence
24Claude Sonnet 5.561.5%estimated ± 5.9 pp, low confidence
25DeepSeek V4.1 Flash61.5%estimated ± 5.9 pp, low confidence
26DeepSeek V4 Flash 073161.5%estimated ± 5.9 pp, low confidence
27DeepSeek V4 Pro 081361.5%estimated ± 5.9 pp, low confidence
28Gemini 2.5 Pro61.5%estimated ± 5.9 pp, medium confidence
29Gemini 3.1 Pro61.5%estimated ± 5.9 pp, low confidence
30Gemini 3.5 Flash61.5%estimated ± 5.9 pp, medium confidence
31Gemini 3.5 Flash-Lite61.5%estimated ± 5.9 pp, medium confidence
32Gemini 3.6 Flash61.5%estimated ± 5.9 pp, low confidence
33Gemini 3.7 Flash61.5%estimated ± 5.9 pp, low confidence
34Gemini 3.8 Flash61.5%estimated ± 5.9 pp, low confidence
35Gemini 3 Pro61.5%estimated ± 5.9 pp, medium confidence
36Gemini 4 Argon61.5%estimated ± 5.9 pp, low confidence
37Gemma 4 26B A4B61.5%estimated ± 5.9 pp, medium confidence
38Gemma 4 31B61.5%estimated ± 5.9 pp, medium confidence
39GLM-4.761.5%estimated ± 5.9 pp, medium confidence
40GLM-5.161.5%estimated ± 5.9 pp, medium confidence
41GLM-5.261.5%estimated ± 5.9 pp, medium confidence
42GLM-5.361.5%estimated ± 5.9 pp, low confidence
43GLM-5.3-Flash61.5%estimated ± 5.9 pp, low confidence
44GLM-5-Turbo61.5%estimated ± 5.9 pp, medium confidence
45GLM-5V-Turbo61.5%estimated ± 5.9 pp, medium confidence
46GPT-4.161.5%estimated ± 5.9 pp, medium confidence
47GPT-5.161.5%estimated ± 5.9 pp, low confidence
48GPT-5.1-Codex61.5%estimated ± 5.9 pp, medium confidence
49GPT-5.1-Codex-Max61.5%estimated ± 5.9 pp, medium confidence
50GPT-5.261.5%estimated ± 5.9 pp, low confidence
51GPT-5.2-Codex61.5%estimated ± 5.9 pp, low confidence
52GPT-5.3 Codex61.5%estimated ± 5.9 pp, low confidence
53GPT-5.461.5%estimated ± 5.9 pp, low confidence
54GPT-5.4 mini61.5%estimated ± 5.9 pp, medium confidence
55GPT-5.4 nano61.5%estimated ± 5.9 pp, medium confidence
56GPT-5.561.5%estimated ± 5.9 pp, low confidence
57GPT-5.6 Luna61.5%estimated ± 5.9 pp, low confidence
58GPT-5.6 Sol61.5%estimated ± 5.9 pp, low confidence
59GPT-5.6 Terra61.5%estimated ± 5.9 pp, low confidence
60GPT-5 (high)61.5%estimated ± 5.9 pp, medium confidence
61GPT-5 (medium)61.5%estimated ± 5.9 pp, medium confidence
62GPT-6.1 Sol61.5%estimated ± 5.9 pp, low confidence
63GPT-6 Astra61.5%estimated ± 5.9 pp, low confidence
64GPT-6 Luna61.5%estimated ± 5.9 pp, low confidence
65GPT-6 Sol61.5%estimated ± 5.9 pp, low confidence
66Grok 461.5%estimated ± 5.9 pp, medium confidence
67Grok 4.1 Fast (Reasoning)61.5%estimated ± 5.9 pp, medium confidence
68Grok 4.561.5%estimated ± 5.9 pp, low confidence
69Grok 4.661.5%estimated ± 5.9 pp, low confidence
70Grok 4.761.5%estimated ± 5.9 pp, medium confidence
71Grok 4 Fast (Reasoning)61.5%estimated ± 5.9 pp, medium confidence
72Hy361.5%estimated ± 5.9 pp, low confidence
73Hy3 Preview61.5%estimated ± 5.9 pp, medium confidence
74Inkling61.5%estimated ± 5.9 pp, medium confidence
75Inkling-Small61.5%estimated ± 5.9 pp, medium confidence
76Kimi K2.661.5%estimated ± 5.9 pp, low confidence
77Kimi K2.5 (Reasoning)61.5%estimated ± 5.9 pp, medium confidence
78Kimi K2.7 Code61.5%estimated ± 5.9 pp, low confidence
79Kimi K361.5%estimated ± 5.9 pp, low confidence
80Ling 3.0 Flash61.5%estimated ± 5.9 pp, medium confidence
81Ling 3.0 Flash FP861.5%estimated ± 5.9 pp, medium confidence
82Ling 3.0 Flash VL61.5%estimated ± 5.9 pp, medium confidence
83Ling 3.1 Flash61.5%estimated ± 5.9 pp, low confidence
84Mercury 2.561.5%estimated ± 5.9 pp, medium confidence
85MiMo-V2.5-Pro61.5%estimated ± 5.9 pp, low confidence
86MiMo-V2.6-Flash61.5%estimated ± 5.9 pp, medium confidence
87MiMo-V2.6-Pro61.5%estimated ± 5.9 pp, low confidence
88MiMo-V2-Omni61.5%estimated ± 5.9 pp, medium confidence
89MiMo-V2-Pro61.5%estimated ± 5.9 pp, medium confidence
90MiniMax M2.761.5%estimated ± 5.9 pp, medium confidence
91MiniMax M361.5%estimated ± 5.9 pp, low confidence
92Mistral Large 461.5%estimated ± 5.9 pp, low confidence
93Mistral Medium 3.5 128B61.5%estimated ± 5.9 pp, medium confidence
94Muse Glimmer 30B61.5%estimated ± 5.9 pp, low confidence
95Muse Spark61.5%estimated ± 5.9 pp, medium confidence
96Muse Spark 1.161.5%estimated ± 5.9 pp, medium confidence
97Muse Spark 1.261.5%estimated ± 5.9 pp, low confidence
98Muse Spark 1.361.5%estimated ± 5.9 pp, low confidence
99Nemotron 3 Super 100B61.5%estimated ± 5.9 pp, medium confidence
100o361.5%estimated ± 5.9 pp, medium confidence
101Quasar 438B61.5%estimated ± 5.9 pp, medium confidence
102Qwen3.6-27B61.5%estimated ± 5.9 pp, medium confidence
103Qwen3.6-35B-A3B61.5%estimated ± 5.9 pp, medium confidence
104Qwen 3.6 Max (preview)61.5%estimated ± 5.9 pp, low confidence
105Qwen3.7 Max61.5%estimated ± 5.9 pp, low confidence
106Qwen3.7 Plus61.5%estimated ± 5.9 pp, medium confidence
107Qwen3.8-27B61.5%estimated ± 5.9 pp, low confidence
108Qwen3.8-Flash-Next61.5%estimated ± 5.9 pp, low confidence
109Qwen3.8 Max Preview61.5%estimated ± 5.9 pp, low confidence
110Solar Pro 461.5%estimated ± 5.9 pp, medium confidence
111Step 3.7 Flash61.5%estimated ± 5.9 pp, medium confidence
112Step 5 Preview61.5%estimated ± 5.9 pp, low confidence
113Ternary Bonsai 2 27B61.5%estimated ± 5.9 pp, medium confidence
114o161.5%estimated ± 5.9 pp, medium confidence
115Grok 4.361.5%estimated ± 5.9 pp, medium confidence
116Qwen3.5 397B (Reasoning)61.5%estimated ± 5.9 pp, medium confidence
117Gemma 4 12B61.5%estimated ± 5.9 pp, medium confidence
118Solar Open 261.4%estimated ± 5.9 pp, medium confidence
119K-Exaone61.2%estimated ± 5.9 pp, medium confidence
120Kimi K2.561.0%measured
121GLM-560.8%measured
122Qwen3.5-27B60.6%measured
123Agents-A160.2%measured
124Qwen3.5-122B-A10B60.2%measured
125Ling 3.0 Tiny59.7%estimated ± 5.9 pp, medium confidence
126Qwen3.5-35B-A3B59.0%measured
127Agents-A1-4B52.1%measured
128MiniCPM5-2B43.7%measured
129LLaDA2.2-mini35.0%measured
130DeepSeek V3.1 (Reasoning)1.5%estimated ± 5.9 pp, low confidence
131K-EXAONE 2.00.6%estimated ± 5.9 pp, low confidence
132DeepSeek-R10.2%estimated ± 5.9 pp, low confidence
133Gemini 3 Flash0.1%estimated ± 5.9 pp, low confidence
134Grok Code Fast 10.0%estimated ± 5.9 pp, low confidence
135Kimi K20.0%estimated ± 5.9 pp, low confidence
136Command A+0.0%estimated ± 5.9 pp, low confidence
137GPT-OSS 120B0.0%estimated ± 5.9 pp, low confidence
138Celeris-10.0%estimated ± 5.9 pp, low confidence
139Claude 3 Haiku0.0%estimated ± 5.9 pp, low confidence
140Claude 4 Sonnet0.0%estimated ± 5.9 pp, low confidence
141DeepSeek R1 Distill Qwen 32B0.0%estimated ± 5.9 pp, low confidence
142DeepSeek V30.0%estimated ± 5.9 pp, low confidence
143DeepSeek V3 03240.0%estimated ± 5.9 pp, low confidence
144DeepSeek V3.10.0%estimated ± 5.9 pp, low confidence
145DeepSeek V3.20.0%estimated ± 5.9 pp, low confidence
146Exaone 4.0 1.2B0.0%estimated ± 5.9 pp, low confidence
147Gemini 2.5 Flash0.0%estimated ± 5.9 pp, low confidence
148Gemma 3 27B0.0%estimated ± 5.9 pp, low confidence
149Gemma 4 E2B0.0%estimated ± 5.9 pp, low confidence
150Gemma 4 E4B0.0%estimated ± 5.9 pp, low confidence
151GLM-4.5-Air0.0%estimated ± 5.9 pp, low confidence
152GLM-4.60.0%estimated ± 5.9 pp, low confidence
153GPT-4.1 mini0.0%estimated ± 5.9 pp, low confidence
154GPT-4.1 nano0.0%estimated ± 5.9 pp, low confidence
155GPT-4o0.0%estimated ± 5.9 pp, low confidence
156GPT-OSS 20B0.0%estimated ± 5.9 pp, low confidence
157Granite-4.0-350M0.0%estimated ± 5.9 pp, low confidence
158Granite-4.0-H-1B0.0%estimated ± 5.9 pp, low confidence
159Granite-4.0-H-350M0.0%estimated ± 5.9 pp, low confidence
160Granite 4.2 30B0.0%estimated ± 5.9 pp, low confidence
161Granite 4.2 3B0.0%estimated ± 5.9 pp, low confidence
162Granite 4.2 8B0.0%estimated ± 5.9 pp, low confidence
163Grok 4.1 Fast0.0%estimated ± 5.9 pp, low confidence
164LFM2.5-2.6B0.0%estimated ± 5.9 pp, low confidence
165LFM2.5-8B-A1B0.0%estimated ± 5.9 pp, low confidence
166LFM2.5-VL-1.6B-Extract0.0%estimated ± 5.9 pp, low confidence
167Ling 2.6 Flash0.0%estimated ± 5.9 pp, low confidence
168Llama 3.1 405B0.0%estimated ± 5.9 pp, low confidence
169Llama 4 Maverick0.0%estimated ± 5.9 pp, low confidence
170Llama 4 Scout0.0%estimated ± 5.9 pp, low confidence
171MiMo-V2-Flash0.0%estimated ± 5.9 pp, low confidence
172Mistral Large 30.0%estimated ± 5.9 pp, low confidence
173Mistral Medium 30.0%estimated ± 5.9 pp, low confidence
174Mistral Small 40.0%estimated ± 5.9 pp, low confidence
175Mistral Small 4 (Reasoning)0.0%estimated ± 5.9 pp, low confidence
176Nemotron 3.5 Lightning 30B A3B NVFP40.0%estimated ± 5.9 pp, low confidence
177Nemotron 3 Nano 30B0.0%estimated ± 5.9 pp, low confidence
178Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 5.9 pp, low confidence
179North Mini Code0.0%estimated ± 5.9 pp, low confidence
180Nova Pro0.0%estimated ± 5.9 pp, low confidence
181Phi-40.0%estimated ± 5.9 pp, low confidence
182Qwen3 Max0.0%estimated ± 5.9 pp, low confidence
183Qwen3-Omni-30B-A3B-Instruct0.0%estimated ± 5.9 pp, low confidence
184Qwen3-Omni-30B-A3B-Thinking0.0%estimated ± 5.9 pp, low confidence
185Sarvam 105B0.0%estimated ± 5.9 pp, low confidence
186Sarvam 30B0.0%estimated ± 5.9 pp, low confidence
187Solar Pro 20.0%estimated ± 5.9 pp, low confidence
188Solar Pro 30.0%estimated ± 5.9 pp, low confidence
189Trinity-Large-Preview0.0%estimated ± 5.9 pp, low confidence
190Trinity-Large-Thinking0.0%estimated ± 5.9 pp, low confidence
191Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 5.9 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General