benchgap
Agentic · tools

VITA-Bench leaderboard

As of 2026-10-07, the highest measured score on VITA-Bench is 47.9% by Qwen3.7 Max. 71 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1K-EXAONE 2.055.6%estimated ± 10.7 pp, low confidence
2Ornith-1.0-397B54.9%estimated ± 10.7 pp, low confidence
3Muse Spark 1.151.3%estimated ± 9.9 pp, low confidence
4Claude Opus 550.4%estimated ± 9.9 pp, low confidence
5Step 5 Preview50.3%estimated ± 9.9 pp, low confidence
6Kimi K349.8%estimated ± 9.9 pp, low confidence
7Hy4 preview49.6%estimated ± 9.9 pp, low confidence
8Gemini 3.5 Flash49.6%estimated ± 9.9 pp, low confidence
9Claude Opus 4.849.0%estimated ± 9.9 pp, low confidence
10Ornith-1.5-397B48.2%estimated ± 9.9 pp, low confidence
11Inkling-Small48.0%estimated ± 9.9 pp, low confidence
12Qwen3.7 Max47.9%measured
13Beam47.7%estimated ± 9.9 pp, low confidence
14Claude Opus 4.7 (Adaptive)47.1%estimated ± 9.9 pp, low confidence
15GLM-5.246.9%estimated ± 9.9 pp, low confidence
16Ornith-1.0-35B46.9%estimated ± 10.7 pp, low confidence
17Kimi K2.7 Code46.6%estimated ± 9.9 pp, low confidence
18Atria Dawn Preview46.6%estimated ± 9.9 pp, low confidence
19GPT-5.6 Sol46.6%estimated ± 9.9 pp, low confidence
20GPT-6 Astra46.5%estimated ± 9.9 pp, low confidence
21Muse Glimmer 30B46.4%estimated ± 9.9 pp, low confidence
22GPT-5.5 Pro46.3%estimated ± 9.9 pp, low confidence
23GPT-5.546.3%estimated ± 9.9 pp, low confidence
24GPT-5.4 Pro46.2%estimated ± 9.9 pp, low confidence
25Claude Mythos 546.1%estimated ± 9.9 pp, low confidence
26GPT-5.6 Terra46.0%estimated ± 9.9 pp, low confidence
27MiniMax M345.9%estimated ± 9.9 pp, low confidence
28Inkling45.8%estimated ± 9.9 pp, low confidence
29DeepSeek V4 Pro 081345.6%estimated ± 9.9 pp, low confidence
30Qwen3.7 Plus45.6%measured
31Claude Sonnet 545.6%estimated ± 9.9 pp, low confidence
32Claude Opus 4.645.4%estimated ± 9.9 pp, low confidence
33dots3-note Preview45.3%estimated ± 9.9 pp, low confidence
34GPT-5.6 Luna45.3%estimated ± 9.9 pp, low confidence
35Claude Sonnet 4.644.9%estimated ± 10.7 pp, low confidence
36GLM-5.144.9%estimated ± 9.9 pp, low confidence
37GPT-5.444.3%estimated ± 9.9 pp, low confidence
38Qwen3.6 Plus44.3%measured
39Ornith-1.5-35B-A3B44.2%estimated ± 9.9 pp, low confidence
40Qwen3.5 397B43.7%measured
41DeepSeek V4 Flash 073143.7%estimated ± 9.9 pp, low confidence
42Step 3.7 Flash43.4%estimated ± 9.9 pp, low confidence
43Ling 3.0 Flash42.1%estimated ± 9.9 pp, low confidence
44Qwen 3.6 Max (preview)41.6%estimated ± 8.2 pp, low confidence
45MiMo-V2.5-Pro41.0%estimated ± 10.7 pp, low confidence
46Muse Spark41.0%estimated ± 10.7 pp, low confidence
47Ornith-1.0-9B40.3%estimated ± 10.7 pp, low confidence
48Agents-A1-4B40.3%measured
49Solar Pro 440.3%estimated ± 9.9 pp, low confidence
50LFM2.5-2.6B40.1%estimated ± 10.7 pp, low confidence
51MiMo-V2.539.6%estimated ± 10.7 pp, low confidence
52Kimi K2.538.9%estimated ± 8.2 pp, low confidence
53GLM-538.8%estimated ± 8.2 pp, low confidence
54Agents-A138.8%measured
55Solar Open 238.7%estimated ± 9.9 pp, low confidence
56GPT-5.4 mini38.5%estimated ± 9.9 pp, low confidence
57GPT-5.238.4%estimated ± 9.9 pp, low confidence
58Qwen3.6-27B38.4%estimated ± 8.2 pp, low confidence
59GPT-5.4 nano37.7%estimated ± 9.9 pp, low confidence
60Kimi K2.637.6%estimated ± 9.9 pp, low confidence
61Qwen3.5-122B-A10B36.9%estimated ± 9.9 pp, low confidence
62Ornith-1.5-9B36.8%estimated ± 9.9 pp, low confidence
63Qwen3.6-35B-A3B35.6%measured
64Gemini 3.1 Pro35.6%estimated ± 10.7 pp, low confidence
65MiMo-V2-Pro35.6%estimated ± 10.7 pp, low confidence
66LLaDA2.2-mini35.0%estimated ± 10.7 pp, low confidence
67Qwen3.5-27B34.5%estimated ± 9.9 pp, low confidence
68Qwen3.5-35B-A3B34.5%estimated ± 9.9 pp, low confidence
69Kimi K2.5 (Reasoning)34.1%estimated ± 9.9 pp, low confidence
70GLM-5-Turbo33.9%estimated ± 10.7 pp, low confidence
71LLaDA2.2-flash32.6%estimated ± 9.9 pp, low confidence
72GLM-5V-Turbo32.2%estimated ± 10.7 pp, low confidence
73Gemini 3 Flash28.6%estimated ± 10.7 pp, low confidence
74MiniMax M2.728.2%estimated ± 10.7 pp, low confidence
75MiMo-V2-Omni25.6%estimated ± 10.7 pp, low confidence
76Claude Opus 4.523.3%measured
77LongCat-Flash-Lite-Sparse21.7%measured
78DeepSeek V3.218.5%measured
79Claude Sonnet 4.517.0%measured
80GLM-4.715.5%measured
81Nemotron 3 Ultra13.1%estimated ± 9.9 pp, low confidence
82Nemotron 3.5 Lightning 30B A3B NVFP45.2%estimated ± 9.9 pp, low confidence
83Nemotron 3 Super 100B2.5%estimated ± 10.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General