benchgap
Knowledge & reasoning

BioMysteryBench (human-difficult) leaderboard

As of 2026-10-07, the highest measured score on BioMysteryBench (human-difficult) is 56.5% by Gemini 3.8 Flash. 197 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Gemini 3.8 Flash56.5%measured
2Claude Fable 5.151.6%estimated ± 5.9 pp, low confidence
3Claude Fable 551.1%estimated ± 5.9 pp, low confidence
4GPT-6 Astra50.2%estimated ± 5.9 pp, low confidence
5GPT-6.1 Sol50.1%estimated ± 5.9 pp, low confidence
6Claude Opus 5.550.0%measured
7GPT-5.5 Pro49.8%estimated ± 6.8 pp, low confidence
8GPT-5.4 Pro49.7%estimated ± 6.8 pp, low confidence
9Claude Opus 549.4%measured
10GPT-5.6 Sol49.2%estimated ± 5.9 pp, low confidence
11Gemini 3 Pro Deep Think49.1%estimated ± 6.8 pp, low confidence
12GPT-5.548.7%estimated ± 5.9 pp, low confidence
13Gemini 3 Pro47.9%estimated ± 5.9 pp, low confidence
14Claude Haiku 5.547.8%estimated ± 6.8 pp, low confidence
15Gemini 3.1 Pro47.5%estimated ± 5.9 pp, low confidence
16GPT-6 Sol47.4%estimated ± 5.9 pp, low confidence
17Hy4 preview47.2%estimated ± 6.8 pp, low confidence
18GPT-5.3 Codex46.8%estimated ± 5.9 pp, low confidence
19GLM-5.3-Flash46.7%estimated ± 6.8 pp, low confidence
20Muse Spark 1.146.5%estimated ± 5.9 pp, low confidence
21Gemini 3.5 Flash46.4%estimated ± 5.9 pp, low confidence
22Grok 4.546.3%estimated ± 5.9 pp, low confidence
23GPT-5.445.9%estimated ± 5.9 pp, low confidence
24Gemini 3.6 Flash45.6%estimated ± 5.9 pp, low confidence
25Gemini 4 Argon45.6%estimated ± 5.9 pp, low confidence
26Muse Spark45.4%estimated ± 5.9 pp, low confidence
27DeepSeek V4 Pro 081345.2%estimated ± 5.9 pp, low confidence
28Claude Opus 4.7 (Adaptive)45.1%estimated ± 5.9 pp, low confidence
29Claude Opus 4.845.1%estimated ± 5.9 pp, low confidence
30Grok 4.644.8%estimated ± 5.9 pp, low confidence
31Claude Sonnet 5.544.7%measured
32Kimi K344.6%estimated ± 5.9 pp, low confidence
33Grok 4.744.5%estimated ± 5.9 pp, low confidence
34Claude Opus 4.6 (Adaptive)44.3%estimated ± 5.9 pp, low confidence
35GPT-5.6 Terra44.2%estimated ± 5.9 pp, low confidence
36Claude Opus 4.5 Thinking44.1%estimated ± 5.9 pp, low confidence
37DeepSeek V4.1 Flash44.1%estimated ± 5.9 pp, low confidence
38Claude Opus 4.643.8%estimated ± 5.9 pp, low confidence
39Gemini 3 Flash43.8%estimated ± 5.9 pp, low confidence
40Muse Spark 1.243.6%estimated ± 5.9 pp, low confidence
41Gemini 3.7 Flash43.5%measured
42Claude Opus 4.743.3%estimated ± 5.9 pp, low confidence
43GPT-5.243.1%estimated ± 5.9 pp, low confidence
44GPT-6 Luna42.8%estimated ± 5.9 pp, low confidence
45Muse Spark 1.342.7%estimated ± 5.9 pp, low confidence
46GPT-5.6 Luna42.3%estimated ± 5.9 pp, low confidence
47Inkling41.8%estimated ± 5.9 pp, low confidence
48Step 5 Preview41.7%estimated ± 5.9 pp, low confidence
49GPT-5.2-Codex41.5%estimated ± 5.9 pp, low confidence
50Claude Opus 4.541.4%estimated ± 5.9 pp, low confidence
51Grok 441.2%estimated ± 5.9 pp, low confidence
52DeepSeek V4 Flash 073141.1%estimated ± 5.9 pp, low confidence
53GPT-5 (high)41.1%estimated ± 5.9 pp, low confidence
54Claude Sonnet 541.0%estimated ± 5.9 pp, low confidence
55GPT-5.1-Codex40.9%estimated ± 5.9 pp, low confidence
56GPT-5.1-Codex-Max40.9%estimated ± 5.9 pp, low confidence
57Kimi K2.7 Code40.7%estimated ± 5.9 pp, low confidence
58GPT-5 (medium)40.7%estimated ± 5.9 pp, low confidence
59Gemini 2.5 Pro40.5%estimated ± 5.9 pp, low confidence
60Claude Sonnet 4.640.2%estimated ± 5.9 pp, low confidence
61o340.2%estimated ± 5.9 pp, low confidence
62Qwen 3.6 Max (preview)39.8%estimated ± 5.9 pp, low confidence
63GPT-5.139.7%estimated ± 5.9 pp, low confidence
64GPT-5.4 mini39.6%estimated ± 5.9 pp, low confidence
65Kimi K2.538.3%estimated ± 5.9 pp, low confidence
66Kimi K2.5 (Reasoning)38.3%estimated ± 5.9 pp, low confidence
67MiMo-V2.6-Pro38.0%estimated ± 5.9 pp, low confidence
68Grok 4.337.9%estimated ± 5.9 pp, low confidence
69o137.8%estimated ± 5.9 pp, low confidence
70GLM-5.337.5%estimated ± 5.9 pp, low confidence
71Inkling-Small37.1%estimated ± 5.9 pp, low confidence
72Kimi K2.636.7%estimated ± 5.9 pp, low confidence
73Hy336.3%estimated ± 5.9 pp, low confidence
74Apodex 1.1 Mini36.1%estimated ± 5.9 pp, low confidence
75Qwen3.8 Max Preview36.1%estimated ± 5.9 pp, low confidence
76Hy3 Preview36.0%estimated ± 5.9 pp, low confidence
77Qwen3.7 Max35.7%estimated ± 5.9 pp, low confidence
78DeepSeek-R135.3%estimated ± 5.9 pp, low confidence
79Apodex 1.135.3%measured
80Claude 3 Opus35.2%estimated ± 8.1 pp, low confidence
81Claude 4.1 Opus35.2%estimated ± 8.1 pp, low confidence
82DeepSeek R1 Distill Qwen 32B35.2%estimated ± 8.1 pp, low confidence
83Gemini 1.0 Pro35.2%estimated ± 8.1 pp, low confidence
84Gemini 1.5 Pro35.2%estimated ± 8.1 pp, low confidence
85GPT-4 Turbo35.2%estimated ± 8.1 pp, low confidence
86GPT-4o mini35.2%estimated ± 8.1 pp, low confidence
87o1-preview35.2%estimated ± 8.1 pp, low confidence
88o1-pro35.2%estimated ± 8.1 pp, low confidence
89o3-mini35.2%estimated ± 8.1 pp, low confidence
90o3-pro35.2%estimated ± 8.1 pp, low confidence
91Phi-4 Multimodal Instruct35.2%estimated ± 8.1 pp, low confidence
92Qwen2.5 Coder 32B Instruct35.2%estimated ± 8.1 pp, low confidence
93Gemini 3.5 Flash-Lite34.6%estimated ± 5.9 pp, low confidence
94GLM-4.734.5%estimated ± 5.9 pp, low confidence
95GLM-5V-Turbo34.5%estimated ± 5.9 pp, low confidence
96Ling 3.1 Flash34.3%estimated ± 5.9 pp, low confidence
97DeepSeek V3.1 (Reasoning)34.3%estimated ± 5.9 pp, low confidence
98GLM-5-Turbo33.9%estimated ± 5.9 pp, low confidence
99GPT-4.133.4%estimated ± 5.9 pp, low confidence
100Kimi K233.1%estimated ± 5.9 pp, low confidence
101MiMo-V2.6-Flash32.8%estimated ± 5.9 pp, low confidence
102Muse Glimmer 30B32.8%estimated ± 5.9 pp, low confidence
103MiniMax M2.732.7%estimated ± 5.9 pp, low confidence
104MiMo-V2-Pro32.5%estimated ± 5.9 pp, low confidence
105Qwen3.6 Plus32.4%estimated ± 5.9 pp, low confidence
106GLM-532.3%estimated ± 5.9 pp, low confidence
107Gemini 2.5 Flash32.2%estimated ± 5.9 pp, low confidence
108Mistral Large 431.9%estimated ± 5.9 pp, low confidence
109Step 3.7 Flash31.9%estimated ± 5.9 pp, low confidence
110GPT-5.4 nano31.9%estimated ± 5.9 pp, low confidence
111DeepSeek V331.7%estimated ± 5.9 pp, low confidence
112Grok 4.1 Fast (Reasoning)31.4%estimated ± 5.9 pp, low confidence
113Mistral Large 331.3%estimated ± 5.9 pp, low confidence
114Llama 4 Maverick31.2%estimated ± 5.9 pp, low confidence
115Mistral Medium 3.5 128B31.1%estimated ± 5.9 pp, low confidence
116Qwen3.5 397B30.9%estimated ± 5.9 pp, low confidence
117Qwen3.5 397B (Reasoning)30.9%estimated ± 5.9 pp, low confidence
118Qwen3.8-Flash-Next30.9%estimated ± 5.9 pp, low confidence
119Qwen3.5-122B-A10B30.8%estimated ± 5.9 pp, low confidence
120Qwen3 Max30.8%estimated ± 5.9 pp, low confidence
121DeepSeek V3 032430.8%estimated ± 5.9 pp, low confidence
122GLM-5.230.8%estimated ± 5.9 pp, low confidence
123Nemotron 3 Super 100B30.8%estimated ± 5.9 pp, low confidence
124DeepSeek V3.230.5%estimated ± 5.9 pp, low confidence
125GLM-5.130.3%estimated ± 5.9 pp, low confidence
126Grok Code Fast 130.1%estimated ± 5.9 pp, low confidence
127Llama 3.1 405B29.9%estimated ± 5.9 pp, low confidence
128DeepSeek V3.129.8%estimated ± 5.9 pp, low confidence
129Grok 4 Fast (Reasoning)29.5%estimated ± 5.9 pp, low confidence
130Claude 4 Sonnet29.4%estimated ± 5.9 pp, low confidence
131Qwen3.7 Plus29.3%estimated ± 5.9 pp, low confidence
132Trinity-Large-Preview29.3%estimated ± 5.9 pp, low confidence
133Trinity-Large-Thinking29.3%estimated ± 5.9 pp, low confidence
134MiMo-V2.5-Pro29.2%estimated ± 5.9 pp, low confidence
135Mercury 2.528.8%estimated ± 5.9 pp, low confidence
136GPT-OSS 120B28.7%estimated ± 5.9 pp, low confidence
137Mistral Small 428.6%estimated ± 5.9 pp, low confidence
138Mistral Small 4 (Reasoning)28.6%estimated ± 5.9 pp, low confidence
139Nemotron 3 Ultra28.5%estimated ± 5.9 pp, low confidence
140GLM-4.628.3%estimated ± 5.9 pp, low confidence
141Qwen3.5-27B27.7%estimated ± 5.9 pp, low confidence
142GPT-4.1 mini27.3%estimated ± 5.9 pp, low confidence
143Nemotron Ultra 253B27.2%estimated ± 5.9 pp, low confidence
144Qwen3.5-35B-A3B27.2%estimated ± 5.9 pp, low confidence
145Gemma 4 31B27.1%estimated ± 5.9 pp, low confidence
146GPT-4o27.0%estimated ± 5.9 pp, low confidence
147Mistral Large 227.0%estimated ± 5.9 pp, low confidence
148Qwen3.6-27B26.7%estimated ± 5.9 pp, low confidence
149MiMo-V2-Omni26.4%estimated ± 5.9 pp, low confidence
150Gemma 4 26B A4B26.2%estimated ± 5.9 pp, low confidence
151Ultravox v0.6 Llama 3.3 70B26.1%estimated ± 5.9 pp, low confidence
152North Mini Code26.0%estimated ± 5.9 pp, low confidence
153Solar Pro 426.0%estimated ± 5.9 pp, low confidence
154Qwen3.6-35B-A3B25.9%estimated ± 5.9 pp, low confidence
155A.X K225.8%estimated ± 5.9 pp, low confidence
156Solar Pro 325.7%estimated ± 5.9 pp, low confidence
157Mistral Medium 325.5%estimated ± 5.9 pp, low confidence
158Ling 3.0 Flash25.4%estimated ± 5.9 pp, low confidence
159Ling 3.0 Flash FP825.4%estimated ± 5.9 pp, low confidence
160Claude 3 Haiku24.8%estimated ± 5.9 pp, low confidence
161Sarvam 105B24.8%estimated ± 5.9 pp, low confidence
162Nemotron 3 Nano 30B24.5%estimated ± 5.9 pp, low confidence
163Grok 4.1 Fast24.4%estimated ± 5.9 pp, low confidence
164Nova Pro24.1%estimated ± 5.9 pp, low confidence
165MiniMax M323.9%estimated ± 5.9 pp, low confidence
166K-Exaone23.6%estimated ± 5.9 pp, low confidence
167GLM-4.5-Air23.5%estimated ± 5.9 pp, low confidence
168Solar Pro 223.3%estimated ± 5.9 pp, low confidence
169GPT-OSS 20B23.2%estimated ± 5.9 pp, low confidence
170Gemma 4 12B22.7%estimated ± 5.9 pp, low confidence
171Ling 2.6 Flash22.7%estimated ± 5.9 pp, low confidence
172MiMo-V2-Flash22.7%estimated ± 5.9 pp, low confidence
173Qwen3.8-27B22.7%estimated ± 5.9 pp, low confidence
174Quasar 438B22.6%estimated ± 5.9 pp, low confidence
175Llama 4 Scout22.3%estimated ± 5.9 pp, low confidence
176Nemotron 3 Nano Omni 30B A3B22.3%estimated ± 5.9 pp, low confidence
177Qwen3-Omni-30B-A3B-Thinking21.6%estimated ± 5.9 pp, low confidence
178Ling 3.0 Flash VL21.4%estimated ± 5.9 pp, low confidence
179Nemotron 3.5 Lightning 30B A3B NVFP421.4%estimated ± 5.9 pp, low confidence
180Qwen3-Omni-30B-A3B-Instruct21.3%estimated ± 5.9 pp, low confidence
181Phi-421.1%estimated ± 5.9 pp, low confidence
182GPT-4.1 nano20.6%estimated ± 5.9 pp, low confidence
183K-EXAONE 2.020.0%estimated ± 5.9 pp, low confidence
184Gemma 3 27B19.8%estimated ± 5.9 pp, low confidence
185Sarvam 30B19.4%estimated ± 5.9 pp, low confidence
186Granite 4.2 8B17.7%estimated ± 5.9 pp, low confidence
187Celeris-117.4%estimated ± 5.9 pp, low confidence
188Exaone 4.0 32B16.9%estimated ± 5.9 pp, low confidence
189Granite 4.2 30B16.3%estimated ± 5.9 pp, low confidence
190LFM2.5-8B-A1B15.3%estimated ± 5.9 pp, low confidence
191Granite 4.2 3B15.1%estimated ± 5.9 pp, low confidence
192Command A+14.7%estimated ± 5.9 pp, low confidence
193Gemma 4 E4B14.3%estimated ± 5.9 pp, low confidence
194Ling 3.0 Tiny14.1%estimated ± 5.9 pp, low confidence
195MiniCPM5-2B14.0%estimated ± 5.9 pp, low confidence
196Gemma 4 E2B11.4%estimated ± 5.9 pp, low confidence
197LFM2.5-VL-1.6B-Extract10.2%estimated ± 5.9 pp, low confidence
198Granite-4.0-H-1B9.2%estimated ± 5.9 pp, low confidence
199Exaone 4.0 1.2B8.9%estimated ± 5.9 pp, low confidence
200LFM2.5-2.6B8.0%estimated ± 5.9 pp, low confidence
201Granite-4.0-350M7.1%estimated ± 5.9 pp, low confidence
202Granite-4.0-H-350M7.0%estimated ± 5.9 pp, low confidence
203Claude 4.1 Opus Thinking0.0%estimated ± 6.8 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General