benchgap
Vision & documents

CountBench leaderboard

As of 2026-10-10, the highest measured score on CountBench is 97.8% by Qwen3.6-27B. 59 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Mythos 599.4%estimated ± 3.7 pp, low confidence
2Claude Opus 4.7 (Adaptive)98.8%estimated ± 3.7 pp, low confidence
3Claude Opus 4.898.5%estimated ± 3.7 pp, low confidence
4GLM-5.3-Flash98.3%estimated ± 3.7 pp, low confidence
5Gemini 3.7 Flash98.1%estimated ± 3.7 pp, low confidence
6Muse Spark 1.198.1%estimated ± 3.7 pp, low confidence
7Claude Sonnet 598.0%estimated ± 3.7 pp, low confidence
8Qwen3.6-27B97.8%measured
9Sakana Fugu-Ultra97.5%estimated ± 3.7 pp, low confidence
10Step 3.7 Flash97.5%estimated ± 0.9 pp, medium confidence
11Gemini 3.1 Pro97.5%estimated ± 2.2 pp, medium confidence
12Qwen3.8 Max97.4%estimated ± 0.8 pp, low confidence
13Kimi K397.4%estimated ± 0.8 pp, low confidence
14Seed 2.1 Pro97.4%estimated ± 0.8 pp, low confidence
15Qwen3.8-Omni-Flash97.4%estimated ± 0.8 pp, low confidence
16Qwen3.8-Flash-Next97.4%estimated ± 0.8 pp, low confidence
17Qwen3.7 Plus97.4%estimated ± 0.8 pp, low confidence
18Seed 2.1 Turbo97.4%estimated ± 0.8 pp, low confidence
19Qwen3.8-27B97.4%estimated ± 0.8 pp, low confidence
20Qwen3.6 Plus97.3%estimated ± 0.8 pp, medium confidence
21dots3-note Preview97.3%estimated ± 0.8 pp, medium confidence
22Gemini 3 Pro97.3%measured
23Kimi K2.697.3%estimated ± 0.8 pp, medium confidence
24Qwen3.6-35B-A3B97.3%estimated ± 1.6 pp, medium confidence
25Qwen3.5 397B97.2%measured
26Sakana Fugu97.1%estimated ± 3.7 pp, low confidence
27Qwen3.5-122B-A10B97.0%estimated ± 0.8 pp, medium confidence
28Qwen3.5-27B96.9%estimated ± 0.8 pp, medium confidence
29Gemini 3.5 Flash96.8%estimated ± 3.7 pp, low confidence
30MiMo-V2.596.7%estimated ± 1.8 pp, medium confidence
31GPT-5.496.2%estimated ± 2.2 pp, medium confidence
32Inkling96.2%estimated ± 3.7 pp, low confidence
33Muse Spark96.0%estimated ± 2.2 pp, medium confidence
34Inkling-Small95.9%estimated ± 3.7 pp, low confidence
35GPT-5.6 Sol95.8%estimated ± 5.5 pp, low confidence
36GPT-5.595.5%estimated ± 5.5 pp, low confidence
37GPT-5.6 Terra95.4%estimated ± 5.5 pp, low confidence
38GPT-5.6 Luna95.0%estimated ± 5.5 pp, low confidence
39Grok 4.395.0%estimated ± 5.5 pp, low confidence
40Pareto 26.995.0%estimated ± 5.5 pp, low confidence
41Gemma 4 31B94.8%estimated ± 5.5 pp, low confidence
42GPT-5.4 mini94.7%estimated ± 5.5 pp, low confidence
43Kimi K2.5 (Reasoning)94.7%estimated ± 3.7 pp, low confidence
44Claude Sonnet 4.694.6%estimated ± 3.7 pp, low confidence
45Step 5 Preview94.6%estimated ± 5.5 pp, low confidence
46Gemma 4 26B A4B94.3%estimated ± 5.5 pp, low confidence
47Kimi K2.594.1%measured
48Interfaze Beta93.8%estimated ± 5.5 pp, low confidence
49Qwen3.5-35B-A3B93.5%estimated ± 0.8 pp, medium confidence
50Ternary Bonsai 2 27B93.4%estimated ± 3.8 pp, high confidence
51Gemini 3.1 Flash-Lite93.2%estimated ± 3.7 pp, low confidence
52GPT-5.4 nano93.0%estimated ± 5.5 pp, low confidence
53Grok 4.2092.6%estimated ± 2.2 pp, medium confidence
54GPT-5.291.9%measured
55Claude Opus 4.691.9%estimated ± 2.2 pp, medium confidence
56MiniMax M391.0%estimated ± 1.8 pp, medium confidence
57Gemma 4 E4B90.7%estimated ± 5.5 pp, low confidence
58Gemma 4 12B90.6%estimated ± 0.8 pp, medium confidence
59Claude Opus 4.590.6%measured
60Nemotron 3 Nano Omni 30B A3B90.3%estimated ± 2.0 pp, medium confidence
61Qwen2.5-VL-32B90.1%estimated ± 5.5 pp, low confidence
62Gemma 4 E2B89.2%estimated ± 5.5 pp, low confidence
63ZAYA1-VL-8B88.1%measured
64LFM2.5-VL-3B87.3%measured
65Command A+83.7%estimated ± 3.7 pp, low confidence
66North Micro Vision Instruct81.0%estimated ± 3.8 pp, high confidence
67Muse Glimmer 30B77.2%estimated ± 2.5 pp, low confidence
68LFM2.5-VL-450M73.3%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General