benchgap
Knowledge & reasoning

BioMysteryBench (human-solvable) leaderboard

As of 2026-10-07, the highest measured score on BioMysteryBench (human-solvable) is 90.1% by Claude Opus 5. 198 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Opus 590.1%measured
2GPT-5.6 Sol89.5%estimated ± 0.6 pp, low confidence
3GPT-6.1 Sol89.5%estimated ± 0.6 pp, medium confidence
4GPT-6 Astra89.5%estimated ± 0.6 pp, medium confidence
5GPT-6 Sol89.5%estimated ± 0.6 pp, medium confidence
6GPT-5.5 Pro89.5%estimated ± 0.6 pp, medium confidence
7GPT-5.4 Pro89.5%estimated ± 0.6 pp, medium confidence
8GPT-5.6 Terra89.5%estimated ± 0.6 pp, medium confidence
9Claude Fable 5.189.5%estimated ± 0.6 pp, medium confidence
10Claude Fable 589.5%estimated ± 0.6 pp, medium confidence
11Gemini 4 Argon89.5%estimated ± 0.6 pp, medium confidence
12GPT-5.589.5%estimated ± 0.6 pp, medium confidence
13MiMo-V2.6-Pro89.5%estimated ± 0.6 pp, medium confidence
14Gemini 3 Pro Deep Think89.5%estimated ± 0.6 pp, medium confidence
15Muse Spark 1.389.4%estimated ± 0.6 pp, medium confidence
16GPT-5.489.4%estimated ± 0.6 pp, medium confidence
17Kimi K389.4%estimated ± 0.6 pp, medium confidence
18Claude Opus 5.589.3%measured
19Claude Opus 4.889.2%estimated ± 0.6 pp, medium confidence
20GLM-5.289.2%estimated ± 0.6 pp, medium confidence
21Step 5 Preview89.2%estimated ± 0.6 pp, medium confidence
22Claude Sonnet 5.589.2%measured
23GPT-5.6 Luna89.2%estimated ± 0.6 pp, medium confidence
24GPT-6 Luna89.0%estimated ± 0.6 pp, medium confidence
25GLM-5.389.0%estimated ± 0.6 pp, medium confidence
26Claude Haiku 5.588.9%estimated ± 0.6 pp, medium confidence
27Gemini 3.8 Flash88.8%measured
28DeepSeek V4 Pro 081388.8%estimated ± 0.6 pp, medium confidence
29Ling 3.1 Flash88.8%estimated ± 0.6 pp, medium confidence
30Gemini 3.1 Pro88.7%estimated ± 0.6 pp, medium confidence
31Grok 4.788.7%estimated ± 0.6 pp, medium confidence
32Muse Spark 1.288.7%estimated ± 0.6 pp, medium confidence
33Qwen3.8 Max Preview88.7%estimated ± 0.6 pp, medium confidence
34Grok 4.688.5%estimated ± 0.6 pp, medium confidence
35Claude Sonnet 588.4%estimated ± 0.6 pp, medium confidence
36GPT-5.3 Codex88.4%estimated ± 0.6 pp, medium confidence
37Hy4 preview88.4%estimated ± 0.6 pp, medium confidence
38DeepSeek V4 Flash 073188.3%estimated ± 0.6 pp, medium confidence
39GLM-5.3-Flash87.8%estimated ± 0.6 pp, medium confidence
40Grok 4.587.8%estimated ± 0.6 pp, medium confidence
41Muse Spark 1.187.6%estimated ± 0.6 pp, medium confidence
42Gemini 3.7 Flash87.1%measured
43DeepSeek V4.1 Flash87.1%estimated ± 0.6 pp, medium confidence
44Qwen3.7 Max86.4%estimated ± 0.6 pp, low confidence
45Gemini 3.5 Flash86.1%estimated ± 0.6 pp, low confidence
46Claude Opus 4.6 (Adaptive)85.7%estimated ± 0.6 pp, low confidence
47Claude Opus 4.7 (Adaptive)85.1%estimated ± 0.6 pp, low confidence
48MiMo-V2.6-Flash85.1%estimated ± 0.6 pp, low confidence
49GPT-5.284.7%estimated ± 0.6 pp, low confidence
50Muse Spark84.4%estimated ± 0.6 pp, low confidence
51Qwen3.8-Flash-Next84.3%estimated ± 0.6 pp, low confidence
52Gemini 3.6 Flash83.8%estimated ± 0.6 pp, low confidence
53Mistral Large 483.8%estimated ± 0.6 pp, low confidence
54GPT-5.4 mini83.3%estimated ± 0.6 pp, low confidence
55Kimi K2.7 Code83.3%estimated ± 0.6 pp, low confidence
56Quasar 438B82.9%estimated ± 0.6 pp, low confidence
57GPT-5.4 nano82.8%estimated ± 0.6 pp, low confidence
58Gemini 3 Pro82.7%estimated ± 0.6 pp, low confidence
59Qwen3.7 Plus82.7%estimated ± 0.6 pp, low confidence
60A.X K282.6%estimated ± 0.6 pp, low confidence
61GPT-5.2-Codex82.5%estimated ± 0.6 pp, low confidence
62Inkling-Small82.3%estimated ± 0.6 pp, low confidence
63Grok 4.382.2%estimated ± 0.6 pp, low confidence
64Kimi K2.682.2%estimated ± 0.6 pp, low confidence
65GPT-5.1-Codex81.8%estimated ± 0.6 pp, low confidence
66GPT-5.1-Codex-Max81.8%estimated ± 0.6 pp, low confidence
67GPT-5 (high)81.8%estimated ± 0.6 pp, low confidence
68Inkling81.8%estimated ± 0.6 pp, low confidence
69Qwen3.8-27B81.8%estimated ± 0.6 pp, low confidence
70Solar Pro 481.8%estimated ± 0.6 pp, low confidence
71Claude Opus 4.781.8%estimated ± 0.6 pp, low confidence
72GPT-5.181.8%estimated ± 0.6 pp, low confidence
73Hy381.8%estimated ± 0.6 pp, low confidence
74Hy3 Preview81.8%estimated ± 0.6 pp, low confidence
75Apodex 1.181.8%estimated ± 0.6 pp, low confidence
76Apodex 1.1 Mini81.8%estimated ± 0.6 pp, low confidence
77Claude Opus 4.5 Thinking81.8%estimated ± 0.6 pp, low confidence
78GLM-5.181.8%estimated ± 0.6 pp, low confidence
79MiMo-V2.5-Pro81.7%estimated ± 0.6 pp, low confidence
80MiniMax M381.7%estimated ± 0.6 pp, low confidence
81Qwen 3.6 Max (preview)81.7%estimated ± 0.6 pp, low confidence
82Kimi K2.581.7%estimated ± 0.6 pp, low confidence
83Kimi K2.5 (Reasoning)81.7%estimated ± 0.6 pp, low confidence
84Nemotron 3 Super 100B81.7%estimated ± 0.6 pp, low confidence
85Nemotron 3 Ultra81.7%estimated ± 0.6 pp, low confidence
86Grok 4.1 Fast (Reasoning)81.7%estimated ± 0.6 pp, low confidence
87Grok 4 Fast (Reasoning)81.7%estimated ± 0.6 pp, low confidence
88Qwen3.6 Plus81.7%estimated ± 0.6 pp, low confidence
89Claude Opus 4.681.7%estimated ± 0.6 pp, low confidence
90Gemini 2.5 Pro81.7%estimated ± 0.6 pp, low confidence
91Muse Glimmer 30B81.7%estimated ± 0.6 pp, low confidence
92Step 3.7 Flash81.7%estimated ± 0.6 pp, low confidence
93DeepSeek V3.1 (Reasoning)81.7%estimated ± 0.6 pp, low confidence
94GLM-581.7%estimated ± 0.6 pp, low confidence
95Grok 481.7%estimated ± 0.6 pp, low confidence
96Ling 3.0 Flash VL81.7%estimated ± 0.6 pp, low confidence
97Celeris-181.7%estimated ± 0.6 pp, low confidence
98Claude 3 Haiku81.7%estimated ± 0.6 pp, low confidence
99Claude 4.1 Opus Thinking81.7%estimated ± 0.6 pp, low confidence
100Claude 4 Sonnet81.7%estimated ± 0.6 pp, low confidence
101Claude Opus 4.581.7%estimated ± 0.6 pp, low confidence
102Claude Sonnet 4.681.7%estimated ± 0.6 pp, low confidence
103Command A+81.7%estimated ± 0.6 pp, low confidence
104DeepSeek-R181.7%estimated ± 0.6 pp, low confidence
105DeepSeek V381.7%estimated ± 0.6 pp, low confidence
106DeepSeek V3 032481.7%estimated ± 0.6 pp, low confidence
107DeepSeek V3.181.7%estimated ± 0.6 pp, low confidence
108DeepSeek V3.281.7%estimated ± 0.6 pp, low confidence
109Exaone 4.0 1.2B81.7%estimated ± 0.6 pp, low confidence
110Exaone 4.0 32B81.7%estimated ± 0.6 pp, low confidence
111Gemini 2.5 Flash81.7%estimated ± 0.6 pp, low confidence
112Gemini 3.5 Flash-Lite81.7%estimated ± 0.6 pp, low confidence
113Gemini 3 Flash81.7%estimated ± 0.6 pp, low confidence
114Gemma 3 27B81.7%estimated ± 0.6 pp, low confidence
115Gemma 4 12B81.7%estimated ± 0.6 pp, low confidence
116Gemma 4 26B A4B81.7%estimated ± 0.6 pp, low confidence
117Gemma 4 31B81.7%estimated ± 0.6 pp, low confidence
118Gemma 4 E2B81.7%estimated ± 0.6 pp, low confidence
119Gemma 4 E4B81.7%estimated ± 0.6 pp, low confidence
120GLM-4.5-Air81.7%estimated ± 0.6 pp, low confidence
121GLM-4.681.7%estimated ± 0.6 pp, low confidence
122GLM-4.781.7%estimated ± 0.6 pp, low confidence
123GLM-5-Turbo81.7%estimated ± 0.6 pp, low confidence
124GLM-5V-Turbo81.7%estimated ± 0.6 pp, low confidence
125GPT-4.181.7%estimated ± 0.6 pp, low confidence
126GPT-4.1 mini81.7%estimated ± 0.6 pp, low confidence
127GPT-4.1 nano81.7%estimated ± 0.6 pp, low confidence
128GPT-4o81.7%estimated ± 0.6 pp, low confidence
129GPT-5 (medium)81.7%estimated ± 0.6 pp, low confidence
130GPT-OSS 120B81.7%estimated ± 0.6 pp, low confidence
131GPT-OSS 20B81.7%estimated ± 0.6 pp, low confidence
132Granite-4.0-350M81.7%estimated ± 0.6 pp, low confidence
133Granite-4.0-H-1B81.7%estimated ± 0.6 pp, low confidence
134Granite-4.0-H-350M81.7%estimated ± 0.6 pp, low confidence
135Granite 4.2 30B81.7%estimated ± 0.6 pp, low confidence
136Granite 4.2 3B81.7%estimated ± 0.6 pp, low confidence
137Granite 4.2 8B81.7%estimated ± 0.6 pp, low confidence
138Grok 4.1 Fast81.7%estimated ± 0.6 pp, low confidence
139Grok Code Fast 181.7%estimated ± 0.6 pp, low confidence
140K-Exaone81.7%estimated ± 0.6 pp, low confidence
141K-EXAONE 2.081.7%estimated ± 0.6 pp, low confidence
142Kimi K281.7%estimated ± 0.6 pp, low confidence
143LFM2.5-2.6B81.7%estimated ± 0.6 pp, low confidence
144LFM2.5-8B-A1B81.7%estimated ± 0.6 pp, low confidence
145LFM2.5-VL-1.6B-Extract81.7%estimated ± 0.6 pp, low confidence
146Ling 2.6 Flash81.7%estimated ± 0.6 pp, low confidence
147Ling 3.0 Flash81.7%estimated ± 0.6 pp, low confidence
148Ling 3.0 Flash FP881.7%estimated ± 0.6 pp, low confidence
149Ling 3.0 Tiny81.7%estimated ± 0.6 pp, low confidence
150Llama 3.1 405B81.7%estimated ± 0.6 pp, low confidence
151Llama 4 Maverick81.7%estimated ± 0.6 pp, low confidence
152Llama 4 Scout81.7%estimated ± 0.6 pp, low confidence
153Mercury 2.581.7%estimated ± 0.6 pp, low confidence
154MiMo-V2-Flash81.7%estimated ± 0.6 pp, low confidence
155MiMo-V2-Omni81.7%estimated ± 0.6 pp, low confidence
156MiMo-V2-Pro81.7%estimated ± 0.6 pp, low confidence
157MiniCPM5-2B81.7%estimated ± 0.6 pp, low confidence
158MiniMax M2.781.7%estimated ± 0.6 pp, low confidence
159Mistral Large 281.7%estimated ± 0.6 pp, low confidence
160Mistral Large 381.7%estimated ± 0.6 pp, low confidence
161Mistral Medium 381.7%estimated ± 0.6 pp, low confidence
162Mistral Medium 3.5 128B81.7%estimated ± 0.6 pp, low confidence
163Mistral Small 481.7%estimated ± 0.6 pp, low confidence
164Mistral Small 4 (Reasoning)81.7%estimated ± 0.6 pp, low confidence
165Nemotron 3.5 Lightning 30B A3B NVFP481.7%estimated ± 0.6 pp, low confidence
166Nemotron 3 Nano 30B81.7%estimated ± 0.6 pp, low confidence
167Nemotron 3 Nano Omni 30B A3B81.7%estimated ± 0.6 pp, low confidence
168Nemotron Ultra 253B81.7%estimated ± 0.6 pp, low confidence
169North Mini Code81.7%estimated ± 0.6 pp, low confidence
170Nova Pro81.7%estimated ± 0.6 pp, low confidence
171o181.7%estimated ± 0.6 pp, low confidence
172o381.7%estimated ± 0.6 pp, low confidence
173Phi-481.7%estimated ± 0.6 pp, low confidence
174Qwen3.5-122B-A10B81.7%estimated ± 0.6 pp, low confidence
175Qwen3.5-27B81.7%estimated ± 0.6 pp, low confidence
176Qwen3.5-35B-A3B81.7%estimated ± 0.6 pp, low confidence
177Qwen3.5 397B81.7%estimated ± 0.6 pp, low confidence
178Qwen3.5 397B (Reasoning)81.7%estimated ± 0.6 pp, low confidence
179Qwen3.6-27B81.7%estimated ± 0.6 pp, low confidence
180Qwen3.6-35B-A3B81.7%estimated ± 0.6 pp, low confidence
181Qwen3 Max81.7%estimated ± 0.6 pp, low confidence
182Qwen3-Omni-30B-A3B-Instruct81.7%estimated ± 0.6 pp, low confidence
183Qwen3-Omni-30B-A3B-Thinking81.7%estimated ± 0.6 pp, low confidence
184Sarvam 105B81.7%estimated ± 0.6 pp, low confidence
185Sarvam 30B81.7%estimated ± 0.6 pp, low confidence
186Solar Pro 281.7%estimated ± 0.6 pp, low confidence
187Solar Pro 381.7%estimated ± 0.6 pp, low confidence
188Trinity-Large-Preview81.7%estimated ± 0.6 pp, low confidence
189Trinity-Large-Thinking81.7%estimated ± 0.6 pp, low confidence
190Ultravox v0.6 Llama 3.3 70B81.7%estimated ± 0.6 pp, low confidence
191o3-pro49.1%estimated ± 1.0 pp, low confidence
192Claude 4.1 Opus27.9%estimated ± 1.0 pp, low confidence
193o3-mini3.6%estimated ± 1.0 pp, low confidence
194o1-pro3.5%estimated ± 1.0 pp, low confidence
195o1-preview2.1%estimated ± 1.0 pp, low confidence
196Claude 3 Opus0.4%estimated ± 1.0 pp, low confidence
197DeepSeek R1 Distill Qwen 32B0.3%estimated ± 1.0 pp, low confidence
198Gemini 1.5 Pro0.2%estimated ± 1.0 pp, low confidence
199GPT-4 Turbo0.1%estimated ± 1.0 pp, low confidence
200Qwen2.5 Coder 32B Instruct0.1%estimated ± 1.0 pp, low confidence
201GPT-4o mini0.1%estimated ± 1.0 pp, low confidence
202Phi-4 Multimodal Instruct0.0%estimated ± 1.0 pp, low confidence
203Gemini 1.0 Pro0.0%estimated ± 1.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General