benchgap
Vision & documents

V* leaderboard

As of 2026-10-07, the highest measured score on V* is 96.9% by Kimi K2.6. 16 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Kimi K3100.0%estimated ± 5.1 pp, low confidence
2Qwen3.8 Max100.0%estimated ± 5.1 pp, low confidence
3Seed 2.1 Pro99.7%estimated ± 5.1 pp, low confidence
4Qwen3.8-Omni-Flash98.8%estimated ± 5.1 pp, low confidence
5Qwen3.8-Flash-Next97.4%estimated ± 5.1 pp, low confidence
6Qwen3.7 Plus97.0%estimated ± 5.1 pp, low confidence
7Kimi K2.696.9%measured
8Qwen3.6 Plus96.9%measured
9Seed 2.1 Turbo96.8%estimated ± 5.1 pp, low confidence
10Qwen3.8-27B96.7%estimated ± 5.1 pp, low confidence
11Qwen3.5 397B95.8%measured
12Step 3.7 Flash95.3%measured
13Qwen3.6-27B94.7%measured
14Qwen3.5-27B93.7%measured
15dots3-note Preview93.6%estimated ± 5.1 pp, medium confidence
16Qwen3.6-35B-A3B93.3%estimated ± 1.4 pp, medium confidence
17Qwen3.5-122B-A10B93.2%measured
18Qwen3.5-35B-A3B92.7%measured
19Command A+89.1%estimated ± 1.4 pp, low confidence
20Gemini 3 Pro88.0%measured
21Nemotron 3 Nano Omni 30B A3B86.2%estimated ± 1.4 pp, low confidence
22Gemma 4 12B80.0%estimated ± 5.1 pp, medium confidence
23GPT-5.275.9%measured
24LFM2.5-VL-3B68.2%estimated ± 1.4 pp, low confidence
25Claude Opus 4.567.0%measured
26ZAYA1-VL-8B66.0%estimated ± 1.4 pp, low confidence
27LFM2.5-VL-450M51.8%estimated ± 1.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General