benchgap
Agentic · tools

BFCL v4 leaderboard

As of 2026-10-07, the highest measured score on BFCL v4 is 88.5% by BTL-3. 187 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1K-EXAONE 2.0100.0%estimated ± 8.6 pp, low confidence
2Ornith-1.0-397B100.0%estimated ± 8.6 pp, low confidence
3Ornith-1.5-397B100.0%estimated ± 8.6 pp, low confidence
4dots3-note Preview93.7%estimated ± 8.6 pp, low confidence
5Ornith-1.5-35B-A3B91.2%estimated ± 8.6 pp, low confidence
6BTL-388.5%measured
7Claude Opus 4.685.2%estimated ± 8.6 pp, low confidence
8Ornith-1.0-35B83.5%estimated ± 8.6 pp, low confidence
9Claude Opus 5.578.6%estimated ± 2.6 pp, medium confidence
10Claude Sonnet 5.578.5%estimated ± 2.6 pp, medium confidence
11Claude Fable 5.178.3%estimated ± 2.6 pp, medium confidence
12Claude Opus 578.2%estimated ± 2.6 pp, medium confidence
13Grok 4.778.2%estimated ± 2.6 pp, medium confidence
14MiMo-V2.6-Pro78.1%estimated ± 2.6 pp, medium confidence
15Muse Spark 1.378.1%estimated ± 2.6 pp, medium confidence
16Qwen3.8 Max Preview78.1%estimated ± 2.6 pp, medium confidence
17GLM-5.378.0%estimated ± 2.6 pp, medium confidence
18Qwen3.8-Flash-Next77.9%estimated ± 2.6 pp, medium confidence
19Gemini 4 Argon77.9%estimated ± 2.6 pp, medium confidence
20Grok 4.677.9%estimated ± 2.6 pp, medium confidence
21Ling 3.1 Flash77.9%estimated ± 2.6 pp, medium confidence
22Claude Fable 577.9%estimated ± 2.6 pp, medium confidence
23GPT-5.6 Sol77.9%estimated ± 2.6 pp, medium confidence
24MiMo-V2.6-Flash77.9%estimated ± 2.6 pp, medium confidence
25Claude Sonnet 4.677.9%estimated ± 8.6 pp, low confidence
26DeepSeek V4.1 Flash77.8%estimated ± 2.6 pp, medium confidence
27DeepSeek V4 Pro 081377.8%estimated ± 2.6 pp, medium confidence
28Step 5 Preview77.8%estimated ± 2.6 pp, medium confidence
29GPT-6.1 Sol77.8%estimated ± 2.6 pp, medium confidence
30GPT-6 Astra77.6%estimated ± 2.6 pp, medium confidence
31Kimi K377.6%estimated ± 2.6 pp, medium confidence
32GPT-6 Sol77.5%estimated ± 2.6 pp, medium confidence
33Muse Spark 1.277.4%estimated ± 2.6 pp, medium confidence
34Claude Sonnet 577.3%estimated ± 2.6 pp, medium confidence
35GPT-5.6 Luna77.3%estimated ± 2.6 pp, medium confidence
36Claude Opus 4.877.3%estimated ± 2.6 pp, medium confidence
37GPT-5.6 Terra77.3%estimated ± 2.6 pp, medium confidence
38DeepSeek V4 Flash 073177.2%estimated ± 2.6 pp, medium confidence
39GPT-6 Luna77.2%estimated ± 2.6 pp, medium confidence
40Gemini 3.8 Flash77.2%estimated ± 2.6 pp, medium confidence
41Mistral Large 477.1%estimated ± 2.6 pp, medium confidence
42Qwen3.8-27B77.1%estimated ± 2.6 pp, medium confidence
43Gemini 3.7 Flash77.0%estimated ± 2.6 pp, medium confidence
44Atria Dawn Preview77.0%measured
45Grok 4.577.0%estimated ± 2.6 pp, medium confidence
46GLM-5.276.9%estimated ± 2.6 pp, medium confidence
47Claude Opus 4.7 (Adaptive)76.8%estimated ± 2.6 pp, medium confidence
48GPT-5.576.8%estimated ± 2.6 pp, medium confidence
49Gemini 3.5 Flash76.8%estimated ± 2.6 pp, medium confidence
50Gemini 3.6 Flash76.4%estimated ± 2.6 pp, medium confidence
51GPT-5.476.2%estimated ± 2.6 pp, medium confidence
52MiniMax M376.2%estimated ± 2.6 pp, medium confidence
53Hy3 Preview76.0%estimated ± 2.6 pp, medium confidence
54Muse Spark 1.176.0%estimated ± 2.6 pp, medium confidence
55Apodex 1.175.9%estimated ± 2.6 pp, medium confidence
56Apodex 1.1 Mini75.9%estimated ± 2.6 pp, medium confidence
57Quasar 438B75.8%estimated ± 2.6 pp, medium confidence
58Ling 3.0 Flash VL75.6%estimated ± 2.6 pp, medium confidence
59Nemotron 3 Ultra75.6%estimated ± 2.6 pp, medium confidence
60Inkling-Small75.3%estimated ± 2.6 pp, high confidence
61MiMo-V2.5-Pro75.3%estimated ± 2.6 pp, high confidence
62GLM-5.175.3%estimated ± 2.6 pp, high confidence
63Solar Pro 475.2%estimated ± 2.6 pp, high confidence
64Qwen3.7 Max75.0%measured
65Grok 4.375.0%estimated ± 2.6 pp, high confidence
66Inkling74.9%estimated ± 2.6 pp, high confidence
67Hy374.8%estimated ± 2.6 pp, high confidence
68Kimi K2.674.5%estimated ± 2.6 pp, high confidence
69Kimi K2.7 Code74.5%estimated ± 2.6 pp, high confidence
70GLM-4.774.3%estimated ± 2.6 pp, high confidence
71GPT-5.4 mini74.3%estimated ± 2.6 pp, high confidence
72Step 3.7 Flash74.3%estimated ± 2.6 pp, high confidence
73MiniMax M2.774.2%estimated ± 2.6 pp, high confidence
74Ornith-1.5-9B74.2%estimated ± 8.6 pp, low confidence
75Muse Spark74.1%estimated ± 2.6 pp, high confidence
76Qwen3.6 Plus74.0%estimated ± 2.6 pp, high confidence
77Qwen3.6-27B73.9%estimated ± 2.6 pp, high confidence
78Gemini 3.5 Flash-Lite73.9%estimated ± 2.6 pp, high confidence
79A.X K273.6%estimated ± 2.6 pp, high confidence
80BTL-473.5%measured
81GPT-5.4 nano73.5%estimated ± 2.6 pp, high confidence
82Ling 3.0 Flash FP873.4%estimated ± 2.6 pp, high confidence
83GPT-5 (high)73.0%estimated ± 2.6 pp, high confidence
84Ling 3.0 Flash73.0%measured
85Qwen3.7 Plus72.9%measured
86Qwen3.6-35B-A3B72.7%estimated ± 2.6 pp, high confidence
87Kimi K2.571.8%estimated ± 2.6 pp, high confidence
88Kimi K2.5 (Reasoning)71.8%estimated ± 2.6 pp, high confidence
89Qwen 3.6 Max (preview)71.5%estimated ± 9.8 pp, low confidence
90GPT-5.171.5%estimated ± 2.6 pp, high confidence
91Qwen3.5-122B-A10B71.2%estimated ± 2.6 pp, high confidence
92Pokee-Isaac 28B70.9%measured
93Gemini 3.1 Pro70.7%estimated ± 2.6 pp, high confidence
94Muse Glimmer 30B70.6%estimated ± 2.6 pp, high confidence
95Mistral Medium 3.5 128B70.0%estimated ± 2.6 pp, high confidence
96Qwen3.5-27B69.4%estimated ± 9.8 pp, low confidence
97Grok 4.1 Fast (Reasoning)68.8%estimated ± 9.8 pp, low confidence
98Claude Opus 4.6 (Adaptive)67.5%estimated ± 9.8 pp, low confidence
99GPT-5.2-Codex67.5%estimated ± 9.8 pp, low confidence
100MiniCPM5-2B66.6%measured
101Nemotron 3.5 Lightning 30B A3B NVFP465.5%estimated ± 2.6 pp, high confidence
102MiMo-V2-Flash65.1%estimated ± 2.6 pp, high confidence
103Claude Opus 4.5 Thinking64.9%estimated ± 9.8 pp, low confidence
104Qwen3.5-35B-A3B64.6%estimated ± 9.8 pp, low confidence
105Gemma 4 31B64.6%estimated ± 2.6 pp, high confidence
106Ornith-1.0-9B64.6%estimated ± 8.6 pp, low confidence
107GPT-OSS 120B64.0%estimated ± 2.6 pp, high confidence
108Gemini 3 Pro62.6%estimated ± 9.8 pp, low confidence
109MiMo-V2.562.3%estimated ± 8.6 pp, low confidence
110GPT-5 (medium)62.0%estimated ± 9.8 pp, low confidence
111GPT-5.3 Codex61.6%estimated ± 9.8 pp, low confidence
112Granite 4.2 30B61.4%measured
113Gemma 4 26B A4B61.0%estimated ± 2.6 pp, high confidence
114LLaDA2.2-flash60.8%measured
115Ling 3.0 Tiny60.7%estimated ± 2.6 pp, high confidence
116GPT-5.260.5%estimated ± 9.8 pp, low confidence
117Qwen3.5 397B (Reasoning)59.7%estimated ± 9.8 pp, low confidence
118GPT-5.1-Codex59.0%estimated ± 9.8 pp, low confidence
119GPT-5.1-Codex-Max59.0%estimated ± 9.8 pp, low confidence
120o357.2%estimated ± 9.8 pp, low confidence
121LFM2.5-2.6B56.9%measured
122Ternary Bonsai 2 27B56.8%estimated ± 9.8 pp, low confidence
123Command A+56.2%estimated ± 2.6 pp, high confidence
124Agents-A1-4B55.3%estimated ± 9.8 pp, low confidence
125Claude Opus 4.554.6%estimated ± 8.6 pp, low confidence
126GLM-4.654.5%estimated ± 9.8 pp, low confidence
127Celeris-153.8%estimated ± 2.6 pp, high confidence
128DeepSeek V353.8%estimated ± 2.6 pp, high confidence
129DeepSeek V3 032453.8%estimated ± 2.6 pp, high confidence
130Gemini 2.5 Pro53.8%estimated ± 2.6 pp, high confidence
131Gemma 3 27B53.8%estimated ± 2.6 pp, high confidence
132Gemma 4 12B53.8%estimated ± 2.6 pp, high confidence
133Gemma 4 E2B53.8%estimated ± 2.6 pp, high confidence
134Gemma 4 E4B53.8%estimated ± 2.6 pp, high confidence
135GPT-4.1 mini53.8%estimated ± 2.6 pp, high confidence
136GPT-4.1 nano53.8%estimated ± 2.6 pp, high confidence
137GPT-4o53.8%estimated ± 2.6 pp, high confidence
138GPT-4o mini53.8%estimated ± 2.6 pp, high confidence
139GPT-OSS 20B53.8%estimated ± 2.6 pp, high confidence
140K-Exaone53.8%estimated ± 2.6 pp, high confidence
141Ling 2.6 Flash53.8%estimated ± 2.6 pp, high confidence
142Llama 4 Maverick53.8%estimated ± 2.6 pp, high confidence
143Llama 4 Scout53.8%estimated ± 2.6 pp, high confidence
144Mercury 2.553.8%estimated ± 2.6 pp, high confidence
145Mistral Large 353.8%estimated ± 2.6 pp, high confidence
146Mistral Small 453.8%estimated ± 2.6 pp, high confidence
147Mistral Small 4 (Reasoning)53.8%estimated ± 2.6 pp, high confidence
148Nemotron 3 Nano 30B53.8%estimated ± 2.6 pp, high confidence
149Nemotron 3 Nano Omni 30B A3B53.8%estimated ± 2.6 pp, high confidence
150Nemotron 3 Super 100B53.8%estimated ± 2.6 pp, high confidence
151North Mini Code53.8%estimated ± 2.6 pp, high confidence
152Solar Pro 353.8%estimated ± 2.6 pp, high confidence
153Trinity-Large-Preview53.8%estimated ± 2.6 pp, high confidence
154Trinity-Large-Thinking53.8%estimated ± 2.6 pp, high confidence
155Ultravox v0.6 Llama 3.3 70B53.8%estimated ± 2.6 pp, high confidence
156Grok Code Fast 153.7%estimated ± 9.8 pp, low confidence
157Grok 453.2%estimated ± 9.8 pp, low confidence
158Qwen3 Max52.9%estimated ± 9.8 pp, low confidence
159Claude Opus 4.752.7%estimated ± 9.8 pp, low confidence
160Granite 4.2 3B52.4%measured
161Granite 4.2 8B52.4%measured
162Claude 4.1 Opus Thinking51.3%estimated ± 9.8 pp, low confidence
163LFM2.5-8B-A1B49.7%measured
164MiMo-V2-Pro49.6%estimated ± 8.6 pp, low confidence
165GLM-549.3%estimated ± 8.6 pp, low confidence
166Grok 4 Fast (Reasoning)49.0%estimated ± 9.8 pp, low confidence
167Grok 4.1 Fast48.3%estimated ± 9.8 pp, low confidence
168o148.0%estimated ± 9.8 pp, low confidence
169LLaDA2.2-mini47.7%measured
170Kimi K247.6%estimated ± 9.8 pp, low confidence
171Qwen3.5 397B46.7%estimated ± 8.6 pp, low confidence
172Claude 4 Sonnet46.0%estimated ± 9.8 pp, low confidence
173Mellum2-12B-A2.5B-Thinking45.6%measured
174GPT-4.145.5%estimated ± 9.8 pp, low confidence
175Sarvam 105B45.5%estimated ± 9.8 pp, low confidence
176GLM-4.5-Air45.5%estimated ± 9.8 pp, low confidence
177DeepSeek V3.1 (Reasoning)45.1%estimated ± 9.8 pp, low confidence
178DeepSeek-R145.0%estimated ± 9.8 pp, low confidence
179DeepSeek V3.145.0%estimated ± 9.8 pp, low confidence
180Sarvam 30B45.0%estimated ± 9.8 pp, low confidence
181Solar Pro 245.0%estimated ± 9.8 pp, low confidence
182Mistral Large 245.0%estimated ± 9.8 pp, low confidence
183o3-mini44.9%estimated ± 9.8 pp, low confidence
184Mistral Medium 344.9%estimated ± 9.8 pp, low confidence
185Qwen3-Omni-30B-A3B-Thinking44.9%estimated ± 9.8 pp, low confidence
186Claude 3 Haiku44.9%estimated ± 9.8 pp, low confidence
187Exaone 4.0 1.2B44.9%estimated ± 9.8 pp, low confidence
188Granite-4.0-H-1B44.9%estimated ± 9.8 pp, low confidence
189Llama 3.1 405B44.9%estimated ± 9.8 pp, low confidence
190Qwen3-Omni-30B-A3B-Instruct44.9%estimated ± 9.8 pp, low confidence
191Gemini 2.5 Flash44.9%estimated ± 9.8 pp, low confidence
192Granite-4.0-H-350M44.9%estimated ± 9.8 pp, low confidence
193Nova Pro44.9%estimated ± 9.8 pp, low confidence
194Granite-4.0-350M44.9%estimated ± 9.8 pp, low confidence
195Nemotron Ultra 253B44.9%estimated ± 9.8 pp, low confidence
196Exaone 4.0 32B44.9%estimated ± 9.8 pp, low confidence
197LFM2.5-VL-1.6B-Extract44.9%estimated ± 9.8 pp, low confidence
198Phi-444.9%estimated ± 9.8 pp, low confidence
199Mellum2-12B-A2.5B-Instruct44.2%measured
200GLM-5-Turbo43.9%estimated ± 8.6 pp, low confidence
201ZAYA1-8B39.2%measured
202GLM-5V-Turbo38.2%estimated ± 8.6 pp, low confidence
203LFM2.5-VL-3B32.5%measured
204Gemini 3 Flash25.2%estimated ± 8.6 pp, low confidence
205MiniCPM5-1B25.2%measured
206LFM2.5-VL-450M21.1%measured
207LFM2.5-230M21.0%measured
208MiMo-V2-Omni13.9%estimated ± 8.6 pp, low confidence
209DeepSeek V3.20.0%estimated ± 8.6 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General