benchgap
Vision & documents

RefCOCO (avg) leaderboard

As of 2026-10-07, the highest measured score on RefCOCO (avg) is 92.5% by Qwen3.6-27B. 6 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.6 Plus92.6%estimated ± 1.8 pp, low confidence
2Qwen3.6-27B92.5%measured
3Qwen3.5-122B-A10B92.4%estimated ± 1.8 pp, low confidence
4Qwen3.5-27B92.2%estimated ± 1.8 pp, medium confidence
5Qwen3.5-35B-A3B92.1%estimated ± 1.8 pp, medium confidence
6Qwen3.6-35B-A3B92.0%measured
7Command A+91.4%estimated ± 1.8 pp, medium confidence
8Nemotron 3 Nano Omni 30B A3B90.5%measured
9LFM2.5-VL-3B87.9%measured
10ZAYA1-VL-8B84.3%measured
11Interfaze Beta82.1%measured
12LFM2.5-VL-450M80.3%estimated ± 1.8 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General