benchgap
Vision & documents

BabyVision leaderboard

As of 2026-10-07, the highest measured score on BabyVision is 82.0% by Qwen3.8 Max. 14 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.8 Max82.0%measured
2Kimi K379.9%estimated ± 4.3 pp, medium confidence
3Muse Spark 1.176.3%measured
4Seed 2.1 Pro73.7%measured
5Qwen3.8-Omni-Flash69.7%estimated ± 4.3 pp, medium confidence
6Qwen3.8-27B65.7%measured
7Qwen3.8-Flash-Next64.7%estimated ± 4.3 pp, medium confidence
8Qwen3.7 Plus63.5%estimated ± 4.3 pp, medium confidence
9Seed 2.1 Turbo62.9%measured
10Qwen3.5 397B56.5%estimated ± 4.3 pp, medium confidence
11Qwen3.6 Plus54.0%estimated ± 4.3 pp, medium confidence
12GLM-5.3-Flash53.4%measured
13Kimi K2.651.6%estimated ± 4.3 pp, low confidence
14dots3-note Preview50.0%measured
15Gemini 3 Pro48.3%estimated ± 4.3 pp, low confidence
16Qwen3.5-122B-A10B46.6%estimated ± 4.3 pp, low confidence
17Qwen3.5-27B45.8%estimated ± 4.3 pp, low confidence
18Qwen3.5-35B-A3B37.2%estimated ± 4.3 pp, low confidence
19GPT-5.233.5%estimated ± 4.3 pp, low confidence
20Gemma 4 12B19.9%estimated ± 4.3 pp, low confidence
21Claude Opus 4.50.0%estimated ± 4.3 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General