benchgap
Multilingual

INCLUDE leaderboard

As of 2026-10-10, the highest measured score on INCLUDE is 89.8% by Claude Opus 5. 51 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 5.590.3%estimated ± 2.0 pp, low confidence
2Claude Mythos 589.8%estimated ± 2.0 pp, low confidence
3Claude Opus 589.8%measured
4Claude Sonnet 5.589.3%estimated ± 2.0 pp, low confidence
5Claude Fable 5.189.0%estimated ± 2.0 pp, medium confidence
6Claude Opus 4.887.6%measured
7Hy4 preview87.2%estimated ± 2.0 pp, medium confidence
8Qwen3.8-Flash-Next86.7%estimated ± 2.0 pp, medium confidence
9Qwen3.8-Omni-Flash86.6%estimated ± 2.0 pp, medium confidence
10Composer 2.586.4%estimated ± 2.0 pp, medium confidence
11Ornith-1.5-397B86.3%estimated ± 2.0 pp, medium confidence
12Claude Opus 4.586.2%measured
13Qwen3.7 Max86.2%measured
14Ornith-1.0-397B86.1%estimated ± 2.0 pp, medium confidence
15Laguna S 2.186.0%estimated ± 2.0 pp, medium confidence
16Claude Sonnet 586.0%estimated ± 2.0 pp, medium confidence
17Grok 4.585.9%estimated ± 2.0 pp, medium confidence
18Beam85.9%estimated ± 2.0 pp, medium confidence
19SWE-1.785.8%estimated ± 2.0 pp, medium confidence
20Qwen3.5 397B85.6%measured
21Kimi K2.685.5%estimated ± 2.0 pp, medium confidence
22MiniMax M2.785.5%estimated ± 2.0 pp, medium confidence
23DeepSeek V4 Pro 081385.4%estimated ± 2.0 pp, medium confidence
24K-EXAONE 2.085.3%estimated ± 1.8 pp, low confidence
25dots3-note Preview85.2%estimated ± 2.0 pp, medium confidence
26Composer 284.7%estimated ± 2.0 pp, medium confidence
27Qwen3.6 Plus84.7%estimated ± 1.4 pp, low confidence
28DeepSeek V4 Flash 073184.6%estimated ± 2.0 pp, medium confidence
29Ling 3.0 Flash84.3%estimated ± 2.0 pp, medium confidence
30MiMo-V2-Flash84.1%estimated ± 2.0 pp, medium confidence
31Ornith-1.5-35B-A3B84.0%estimated ± 2.0 pp, medium confidence
32Qwen3.6-27B84.0%estimated ± 2.0 pp, medium confidence
33GLM-583.8%estimated ± 1.4 pp, low confidence
34Nemotron 3 Ultra83.7%estimated ± 1.4 pp, low confidence
35Ornith-1.0-35B83.5%estimated ± 2.0 pp, medium confidence
36Kimi K2.583.3%measured
37Qwen3.5-122B-A10B83.2%estimated ± 1.4 pp, low confidence
38Qwen3.5-27B83.2%estimated ± 1.4 pp, low confidence
39Qwen3.7 Plus83.0%measured
40Qwen3.6-35B-A3B82.9%estimated ± 2.0 pp, low confidence
41Qwen3 235B 2507 (Reasoning)82.5%estimated ± 1.4 pp, low confidence
42Qwen3.5-35B-A3B82.5%estimated ± 1.4 pp, low confidence
43Laguna M.181.7%estimated ± 2.0 pp, low confidence
44Laguna XS 2.181.7%estimated ± 2.0 pp, low confidence
45Qwen3 235B 250781.6%estimated ± 1.4 pp, low confidence
46LongCat-Flash-Lite-Sparse80.7%estimated ± 2.0 pp, low confidence
47Laguna XS.280.2%estimated ± 2.0 pp, low confidence
48Ornith-1.5-9B79.3%estimated ± 2.0 pp, low confidence
49Ornith-1.0-9B78.6%estimated ± 2.0 pp, low confidence
50GPT-4.177.4%estimated ± 1.4 pp, low confidence
51Kimi K277.3%estimated ± 2.0 pp, low confidence
52DeepSeek V3 032476.0%estimated ± 1.4 pp, low confidence
53Granite 4.2 30B75.8%estimated ± 2.0 pp, low confidence
54Nemotron 3.5 Lightning 30B A3B NVFP474.3%estimated ± 2.0 pp, low confidence
55Granite 4.2 8B72.7%estimated ± 2.0 pp, low confidence
56LLaDA2.2-flash71.1%estimated ± 2.0 pp, low confidence
57GPT-4o69.4%estimated ± 1.4 pp, low confidence
58Phi-460.6%estimated ± 1.4 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General