benchgap
Multilingual

NOVA-63 leaderboard

As of 2026-10-07, the highest measured score on NOVA-63 is 59.1% by Qwen3.5 397B. 47 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.564.5%estimated ± 1.6 pp, low confidence
2Claude Mythos 563.9%estimated ± 1.6 pp, low confidence
3Claude Sonnet 5.563.2%estimated ± 1.6 pp, low confidence
4Claude Opus 562.9%estimated ± 1.6 pp, low confidence
5Claude Fable 5.162.8%estimated ± 1.6 pp, low confidence
6Claude Opus 4.861.0%estimated ± 1.6 pp, low confidence
7Hy4 preview60.4%estimated ± 1.6 pp, low confidence
8Qwen3.8-Flash-Next59.6%estimated ± 1.6 pp, low confidence
9Qwen3.8-Omni-Flash59.4%estimated ± 1.6 pp, low confidence
10Composer 2.559.2%estimated ± 1.6 pp, low confidence
11Qwen3.5 397B59.1%measured
12Ornith-1.5-397B59.1%estimated ± 1.6 pp, low confidence
13Qwen3.7 Max59.0%measured
14Qwen3.7 Plus58.8%measured
15Ornith-1.0-397B58.8%estimated ± 1.6 pp, low confidence
16Laguna S 2.158.6%estimated ± 1.6 pp, low confidence
17Claude Sonnet 558.5%estimated ± 1.6 pp, low confidence
18Grok 4.558.4%estimated ± 1.6 pp, low confidence
19Beam58.4%estimated ± 1.6 pp, low confidence
20SWE-1.758.3%estimated ± 1.6 pp, low confidence
21Qwen3.6 Plus57.9%measured
22Kimi K2.657.9%estimated ± 1.6 pp, low confidence
23MiniMax M2.757.8%estimated ± 1.6 pp, low confidence
24DeepSeek V4 Pro 081357.7%estimated ± 1.6 pp, low confidence
25dots3-note Preview57.4%estimated ± 1.6 pp, low confidence
26Claude Opus 4.556.7%measured
27Nemotron 3 Ultra56.7%estimated ± 1.3 pp, low confidence
28Composer 256.6%estimated ± 1.6 pp, low confidence
29DeepSeek V4 Flash 073156.4%estimated ± 1.6 pp, low confidence
30Qwen3.5-122B-A10B56.3%estimated ± 1.3 pp, low confidence
31Qwen3.5-27B56.3%estimated ± 1.3 pp, low confidence
32Kimi K2.556.0%measured
33Ling 3.0 Flash56.0%estimated ± 1.6 pp, low confidence
34Qwen3.5-35B-A3B55.7%estimated ± 1.3 pp, low confidence
35Ornith-1.5-35B-A3B55.5%estimated ± 1.6 pp, low confidence
36Qwen3.6-27B55.5%estimated ± 1.6 pp, low confidence
37GLM-555.1%measured
38Qwen3 235B 250754.9%estimated ± 1.3 pp, low confidence
39Ornith-1.0-35B54.6%estimated ± 1.6 pp, low confidence
40Qwen3.6-35B-A3B53.6%estimated ± 1.6 pp, low confidence
41Laguna M.151.6%estimated ± 1.6 pp, low confidence
42Laguna XS 2.151.6%estimated ± 1.6 pp, low confidence
43GPT-4.151.5%estimated ± 1.3 pp, low confidence
44DeepSeek V3 032450.3%estimated ± 1.3 pp, low confidence
45LongCat-Flash-Lite-Sparse49.7%estimated ± 1.6 pp, low confidence
46Laguna XS.248.8%estimated ± 1.6 pp, low confidence
47Ornith-1.5-9B47.0%estimated ± 1.6 pp, low confidence
48Ornith-1.0-9B45.7%estimated ± 1.6 pp, low confidence
49GPT-4o45.1%estimated ± 1.3 pp, low confidence
50Granite 4.2 30B39.5%estimated ± 1.6 pp, low confidence
51Phi-438.4%estimated ± 1.3 pp, low confidence
52Nemotron 3.5 Lightning 30B A3B NVFP435.7%estimated ± 1.6 pp, low confidence
53Granite 4.2 8B31.5%estimated ± 1.6 pp, low confidence
54LLaDA2.2-flash26.8%estimated ± 1.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General