benchgap
Coding

Bug Hunt Bench leaderboard

As of 2026-10-09, the highest measured score on Bug Hunt Bench is 51.3% by Claude Sonnet 5.5. 99 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Claude Sonnet 5.551.3%measured
2GPT-6 Astra45.0%measured
3GPT-6.1 Sol44.3%measured
4Gemini 4 Argon44.3%estimated ± 7.2 pp, low confidence
5Claude Fable 5.143.0%measured
6GPT-5.6 Sol42.0%measured
7Claude Opus 5.541.7%measured
8Step 5 Preview36.2%estimated ± 10.9 pp, low confidence
9Claude Fable 535.9%estimated ± 7.4 pp, medium confidence
10Muse Spark 1.332.2%measured
11DeepSeek V4 Flash 073131.8%estimated ± 8.6 pp, low confidence
12GPT-5.6 Terra31.4%estimated ± 8.6 pp, low confidence
13Grok 4.728.8%measured
14GPT-5.6 Luna27.9%estimated ± 8.6 pp, low confidence
15Claude Opus 4.827.2%estimated ± 7.2 pp, low confidence
16Claude Opus 527.0%measured
17Grok 4.627.0%measured
18Claude Sonnet 526.7%estimated ± 8.6 pp, low confidence
19Qwen3.8 Max25.7%measured
20Claude Opus 4.7 (Adaptive)24.6%estimated ± 9.0 pp, low confidence
21GLM-5.224.5%estimated ± 7.2 pp, low confidence
22Qwen3.8-Flash-Next23.7%estimated ± 9.0 pp, low confidence
23MiMo-V2.6-Flash23.4%estimated ± 10.9 pp, low confidence
24MiMo-V2.6-Pro22.7%measured
25Composer 2.522.2%estimated ± 8.6 pp, low confidence
26Gemini 3.7 Flash22.2%estimated ± 7.4 pp, medium confidence
27Qwen3.8 Max Preview21.8%estimated ± 9.0 pp, low confidence
28DeepSeek V4.1 Flash21.7%measured
29Claude Haiku 5.521.5%measured
30Muse Spark 1.121.1%estimated ± 9.0 pp, low confidence
31Kimi K321.0%measured
32Gemini 3.5 Flash19.4%estimated ± 9.0 pp, low confidence
33Claude Haiku 4.519.0%estimated ± 8.6 pp, low confidence
34GLM-5.319.0%measured
35Hy4 preview18.6%estimated ± 10.9 pp, low confidence
36GPT-6 Luna18.3%measured
37Gemini 3.6 Flash18.2%estimated ± 9.0 pp, low confidence
38Gemini 3.8 Flash18.0%measured
39Muse Spark 1.217.9%estimated ± 7.4 pp, low confidence
40GLM-5.3-Flash17.7%measured
41DeepSeek V4 Pro 081317.6%estimated ± 9.0 pp, low confidence
42Claude Opus 4.717.0%estimated ± 7.2 pp, low confidence
43Qwen3.8-27B15.0%measured
44Qwen3.7 Max14.1%estimated ± 9.0 pp, low confidence
45GPT-5.514.0%estimated ± 7.2 pp, low confidence
46Inkling13.9%estimated ± 7.4 pp, low confidence
47Kimi K2.69.9%estimated ± 9.0 pp, low confidence
48Quasar 438B9.4%estimated ± 9.0 pp, low confidence
49Apodex 1.19.1%estimated ± 9.0 pp, low confidence
50Apodex 1.1 Mini9.1%estimated ± 9.0 pp, low confidence
51Kimi K2.7 Code9.1%estimated ± 9.0 pp, low confidence
52MiMo-V2.5-Pro8.6%estimated ± 9.0 pp, low confidence
53Hy37.5%estimated ± 9.0 pp, low confidence
54Hy3 Preview7.5%estimated ± 9.0 pp, low confidence
55Muse Spark7.4%estimated ± 9.0 pp, low confidence
56MiniMax M37.4%estimated ± 9.0 pp, low confidence
57Grok 4.56.8%estimated ± 7.2 pp, low confidence
58GPT-5.4 mini5.8%estimated ± 9.0 pp, low confidence
59GPT-5.4 nano5.8%estimated ± 9.0 pp, low confidence
60Qwen3.7 Plus5.6%estimated ± 9.0 pp, low confidence
61GLM-5.15.6%estimated ± 9.0 pp, low confidence
62Qwen3.6 Plus4.9%estimated ± 9.0 pp, low confidence
63Gemini 3.1 Pro4.9%estimated ± 7.2 pp, low confidence
64Qwen3.6-27B4.5%estimated ± 9.0 pp, low confidence
65Inkling-Small4.1%estimated ± 9.0 pp, low confidence
66MiniMax M2.74.0%estimated ± 9.0 pp, low confidence
67Ling 3.0 Flash3.2%estimated ± 9.0 pp, low confidence
68Ling 3.0 Flash FP83.2%estimated ± 9.0 pp, low confidence
69MiMo-V2-Flash2.9%estimated ± 9.0 pp, low confidence
70GPT-5.12.8%estimated ± 9.0 pp, low confidence
71Gemini 3.5 Flash-Lite2.7%estimated ± 9.0 pp, low confidence
72Nemotron 3 Ultra2.7%estimated ± 9.0 pp, low confidence
73Muse Glimmer 30B2.6%estimated ± 9.0 pp, low confidence
74GPT-5.42.1%estimated ± 7.2 pp, low confidence
75Mistral Medium 3.5 128B2.0%estimated ± 9.0 pp, low confidence
76Kimi K2.52.0%estimated ± 9.0 pp, low confidence
77Kimi K2.5 (Reasoning)2.0%estimated ± 9.0 pp, low confidence
78Qwen3.5-122B-A10B1.7%estimated ± 9.0 pp, low confidence
79GLM-4.71.6%estimated ± 9.0 pp, low confidence
80Gemma 4 31B1.3%estimated ± 9.0 pp, low confidence
81Grok 4.31.1%estimated ± 9.0 pp, low confidence
82Qwen3.6-35B-A3B1.0%estimated ± 9.0 pp, low confidence
83o10.8%estimated ± 9.0 pp, low confidence
84Step 3.7 Flash0.7%estimated ± 9.0 pp, low confidence
85Gemma 4 26B A4B0.7%estimated ± 9.0 pp, low confidence
86GPT-5 (high)0.6%estimated ± 9.0 pp, low confidence
87Nemotron 3 Super 100B0.6%estimated ± 9.0 pp, low confidence
88o1-preview0.3%estimated ± 9.0 pp, low confidence
89Gemini 2.5 Pro0.3%estimated ± 9.0 pp, low confidence
90K-Exaone0.2%estimated ± 9.0 pp, low confidence
91Gemma 4 12B0.2%estimated ± 9.0 pp, low confidence
92GPT-OSS 120B0.2%estimated ± 9.0 pp, low confidence
93Command A+0.1%estimated ± 9.0 pp, low confidence
94Nemotron 3.5 Lightning 30B A3B NVFP40.1%estimated ± 9.0 pp, low confidence
95Mistral Small 40.1%estimated ± 9.0 pp, low confidence
96Mistral Small 4 (Reasoning)0.1%estimated ± 9.0 pp, low confidence
97Trinity-Large-Preview0.1%estimated ± 9.0 pp, low confidence
98Trinity-Large-Thinking0.1%estimated ± 9.0 pp, low confidence
99Ling 2.6 Flash0.1%estimated ± 9.0 pp, low confidence
100Gemini 1.5 Pro0.0%estimated ± 9.0 pp, low confidence
101DeepSeek V30.0%estimated ± 9.0 pp, low confidence
102Granite 4.2 8B0.0%estimated ± 9.0 pp, low confidence
103GPT-4 Turbo0.0%estimated ± 9.0 pp, low confidence
104GPT-OSS 20B0.0%estimated ± 9.0 pp, low confidence
105GPT-4.1 mini0.0%estimated ± 9.0 pp, low confidence
106Mistral Large 30.0%estimated ± 9.0 pp, low confidence
107Claude 3 Opus0.0%estimated ± 9.0 pp, low confidence
108Llama 4 Maverick0.0%estimated ± 9.0 pp, low confidence
109Celeris-10.0%estimated ± 9.0 pp, low confidence
110Nemotron 3 Nano 30B0.0%estimated ± 9.0 pp, low confidence
111Nemotron 3 Nano Omni 30B A3B0.0%estimated ± 9.0 pp, low confidence
112Ultravox v0.6 Llama 3.3 70B0.0%estimated ± 9.0 pp, low confidence
113GPT-4.1 nano0.0%estimated ± 9.0 pp, low confidence
114GPT-4o mini0.0%estimated ± 9.0 pp, low confidence
115Gemma 3 27B0.0%estimated ± 9.0 pp, low confidence
116Gemma 4 E4B0.0%estimated ± 9.0 pp, low confidence
117Llama 4 Scout0.0%estimated ± 9.0 pp, low confidence
118Gemma 4 E2B0.0%estimated ± 9.0 pp, low confidence
119LFM2.5-2.6B0.0%estimated ± 9.0 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General