benchgap
Knowledge & reasoning

BBH leaderboard

As of 2026-10-10, the highest measured score on BBH is 78.8% by Soofi S 30B-A3B. 110 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Celeris-1100.0%estimated ± 3.2 pp, low confidence
2Claude Opus 4.5100.0%estimated ± 3.2 pp, low confidence
3Claude Opus 4.6100.0%estimated ± 3.2 pp, low confidence
4Claude Sonnet 4.6100.0%estimated ± 3.2 pp, low confidence
5DeepSeek-R1100.0%estimated ± 3.2 pp, low confidence
6DeepSeek V3100.0%estimated ± 3.2 pp, low confidence
7DeepSeek V3.1100.0%estimated ± 3.2 pp, low confidence
8DeepSeek V3.1 (Reasoning)100.0%estimated ± 3.2 pp, low confidence
9DeepSeek V4 Flash 0731100.0%estimated ± 3.2 pp, low confidence
10DeepSeek V4 Pro 0813100.0%estimated ± 3.2 pp, low confidence
11Exaone 4.0 32B100.0%estimated ± 3.2 pp, low confidence
12Gemma 4 12B100.0%estimated ± 3.2 pp, low confidence
13Gemma 4 26B A4B100.0%estimated ± 3.2 pp, low confidence
14Gemma 4 31B100.0%estimated ± 3.2 pp, low confidence
15GLM-4.7100.0%estimated ± 3.2 pp, low confidence
16GLM-5100.0%estimated ± 3.2 pp, low confidence
17Granite 4.2 30B100.0%estimated ± 3.2 pp, low confidence
18Granite 4.2 8B100.0%estimated ± 3.2 pp, low confidence
19Grok 3 [Beta]100.0%estimated ± 3.2 pp, low confidence
20K-EXAONE 2.0100.0%estimated ± 3.2 pp, low confidence
21Kimi K2100.0%estimated ± 3.2 pp, low confidence
22Kimi K2.5100.0%estimated ± 3.2 pp, low confidence
23Kimi K2.5 (Reasoning)100.0%estimated ± 3.2 pp, low confidence
24LongCat-Flash-Lite-Sparse100.0%estimated ± 3.2 pp, low confidence
25MAI-Thinking-1100.0%estimated ± 3.2 pp, low confidence
26MiMo-V2-Flash100.0%estimated ± 3.2 pp, low confidence
27MiniCPM5-2B100.0%estimated ± 3.2 pp, low confidence
28Nemotron 3.5 Lightning 30B A3B NVFP4100.0%estimated ± 3.2 pp, low confidence
29Nemotron 3 Nano 30B100.0%estimated ± 3.2 pp, low confidence
30Nemotron 3 Nano Omni 30B A3B100.0%estimated ± 3.2 pp, low confidence
31Nemotron 3 Super 100B100.0%estimated ± 3.2 pp, low confidence
32Nemotron 3 Ultra100.0%estimated ± 3.2 pp, low confidence
33Qwen3 235B 2507100.0%estimated ± 3.2 pp, low confidence
34Qwen3 235B 2507 (Reasoning)100.0%estimated ± 3.2 pp, low confidence
35Qwen3.5-122B-A10B100.0%estimated ± 3.2 pp, low confidence
36Qwen3.5-27B100.0%estimated ± 3.2 pp, low confidence
37Qwen3.5-35B-A3B100.0%estimated ± 3.2 pp, low confidence
38Qwen3.5 397B100.0%estimated ± 3.2 pp, low confidence
39Qwen3.6-27B100.0%estimated ± 3.2 pp, low confidence
40Qwen3.6-35B-A3B100.0%estimated ± 3.2 pp, low confidence
41Qwen3.6 Plus100.0%estimated ± 3.2 pp, low confidence
42Qwen3.7 Max100.0%estimated ± 3.2 pp, low confidence
43Qwen3.7 Plus100.0%estimated ± 3.2 pp, low confidence
44Sarvam 105B100.0%estimated ± 3.2 pp, low confidence
45Sarvam 30B100.0%estimated ± 3.2 pp, low confidence
46Solar Open 2100.0%estimated ± 3.2 pp, low confidence
47Solar Pro 4100.0%estimated ± 3.2 pp, low confidence
48ZAYA1-8B100.0%estimated ± 3.2 pp, low confidence
49Gemma 4 E4B98.9%estimated ± 3.2 pp, low confidence
50Qwen2.5-VL-32B98.2%estimated ± 3.2 pp, low confidence
51ZAYA1-74B-Preview97.4%estimated ± 3.2 pp, low confidence
52Granite 4.2 3B97.1%estimated ± 3.2 pp, low confidence
53Gemma 4 E2B88.1%estimated ± 3.2 pp, low confidence
54Soofi S 30B-A3B78.8%measured
55Claude Mythos 578.7%estimated ± 10.9 pp, low confidence
56Claude Opus 4.7 (Adaptive)78.7%estimated ± 10.9 pp, low confidence
57Claude Opus 4.878.7%estimated ± 10.9 pp, low confidence
58Claude Sonnet 4.578.7%estimated ± 10.9 pp, low confidence
59DeepSeek V4.1 Flash78.7%estimated ± 10.9 pp, low confidence
60Gemini 2.5 Pro78.7%estimated ± 10.9 pp, low confidence
61Gemini 3.5 Flash78.7%estimated ± 10.9 pp, low confidence
62GLM-4.7-Flash78.7%estimated ± 10.9 pp, low confidence
63GLM-5.278.7%estimated ± 10.9 pp, low confidence
64GPT-4.178.7%estimated ± 10.9 pp, low confidence
65GPT-5.278.7%estimated ± 10.9 pp, low confidence
66GPT-5.478.7%estimated ± 10.9 pp, low confidence
67GPT-5.4 mini78.7%estimated ± 10.9 pp, low confidence
68GPT-5.4 nano78.7%estimated ± 10.9 pp, low confidence
69GPT-5.578.7%estimated ± 10.9 pp, low confidence
70GPT-5.6 Luna78.7%estimated ± 10.9 pp, low confidence
71GPT-5.6 Sol78.7%estimated ± 10.9 pp, low confidence
72GPT-5.6 Terra78.7%estimated ± 10.9 pp, low confidence
73GPT-6 Astra78.7%estimated ± 10.9 pp, low confidence
74GPT-OSS 120B78.7%estimated ± 10.9 pp, low confidence
75GPT-OSS 20B78.7%estimated ± 10.9 pp, low confidence
76Grok 3 Mini78.7%estimated ± 10.9 pp, low confidence
77Grok 4.378.7%estimated ± 10.9 pp, low confidence
78Hy3 Preview78.7%estimated ± 10.9 pp, low confidence
79Hy4 preview78.7%estimated ± 10.9 pp, low confidence
80Inkling78.7%estimated ± 10.9 pp, low confidence
81Inkling-Small78.7%estimated ± 10.9 pp, low confidence
82Interfaze Beta78.7%estimated ± 10.9 pp, low confidence
83Kimi K2.678.7%estimated ± 10.9 pp, low confidence
84Kimi K378.7%estimated ± 10.9 pp, low confidence
85Ling 3.0 Flash78.7%estimated ± 10.9 pp, low confidence
86Ling 3.0 Flash FP878.7%estimated ± 10.9 pp, low confidence
87Ministral 3 14B (Reasoning)78.7%estimated ± 10.9 pp, low confidence
88Nemotron Ultra 253B78.7%estimated ± 10.9 pp, low confidence
89o178.7%estimated ± 10.9 pp, low confidence
90o1-pro78.7%estimated ± 10.9 pp, low confidence
91o3-mini78.7%estimated ± 10.9 pp, low confidence
92Ornith-1.5-35B-A3B78.7%estimated ± 10.9 pp, low confidence
93Ornith-1.5-397B78.7%estimated ± 10.9 pp, low confidence
94Ornith-1.5-9B78.7%estimated ± 10.9 pp, low confidence
95Qwen3.8-27B78.7%estimated ± 10.9 pp, low confidence
96Qwen3.8-Flash-Next78.7%estimated ± 10.9 pp, low confidence
97Qwen3.8 Max78.7%estimated ± 10.9 pp, low confidence
98Qwen3.8-Omni-Flash78.7%estimated ± 10.9 pp, low confidence
99Sakana Fugu78.7%estimated ± 10.9 pp, low confidence
100Sakana Fugu-Ultra78.7%estimated ± 10.9 pp, low confidence
101Ternary Bonsai 2 27B78.7%estimated ± 10.9 pp, low confidence
102GPT-4.1 mini78.7%estimated ± 10.9 pp, low confidence
103Claude 3.5 Sonnet78.7%estimated ± 10.9 pp, low confidence
104Ling 2.6 Flash78.7%estimated ± 10.9 pp, low confidence
105Mellum2-12B-A2.5B-Thinking78.7%estimated ± 10.9 pp, low confidence
106Phi-478.7%estimated ± 10.9 pp, low confidence
107GPT-4.1 nano78.7%estimated ± 10.9 pp, low confidence
108Mellum2-12B-A2.5B-Instruct78.7%estimated ± 10.9 pp, low confidence
109MiniCPM5-1B71.9%measured
110Granite-4.0-1B60.4%measured
111Granite-4.0-H-1B59.7%measured
112LFM2.5-230M42.3%estimated ± 3.2 pp, medium confidence
113LFM2.5-350M42.1%estimated ± 3.2 pp, medium confidence
114LFM2.5-VL-450M41.3%estimated ± 3.2 pp, medium confidence
115Granite-4.0-H-350M33.3%measured
116Granite-4.0-350M33.1%measured

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General