benchgap
Instruction following

IFBench leaderboard

As of 2026-10-07, the highest measured score on IFBench is 85.0% by MAI-Thinking-1. 125 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1MAI-Thinking-185.0%measured
2Qwen3.8 Max82.8%measured
3MiniMax M382.3%estimated ± 6.9 pp, low confidence
4Inkling-Small82.2%measured
5Nemotron 3 Ultra81.7%measured
6Qwen3.8-Omni-Flash81.5%measured
7Grok 4.381.3%measured
8Qwen3.8-Flash-Next81.3%measured
9dots3-note Preview80.4%measured
10Solar Open 280.0%measured
11Inkling79.8%measured
12Beam79.7%measured
13Qwen3.8-27B79.5%measured
14Granite 4.2 8B79.3%measured
15Qwen3.7 Max79.1%measured
16Qwen3.7 Plus79.1%measured
17MiMo-V2.5-Pro78.9%estimated ± 6.9 pp, medium confidence
18Granite 4.2 30B77.2%measured
19Mercury 2.577.0%measured
20Muse Glimmer 30B77.0%measured
21GPT-5.2-Codex76.3%estimated ± 6.9 pp, medium confidence
22Gemini 3.5 Flash76.3%measured
23A.X K275.9%measured
24Qwen3.6 Plus75.8%measured
25Gemini 3.1 Pro75.8%estimated ± 6.9 pp, medium confidence
26Qwen 3.6 Max (preview)75.2%estimated ± 6.9 pp, medium confidence
27DeepSeek V4 Pro 081375.1%estimated ± 6.9 pp, medium confidence
28GLM-5.174.9%estimated ± 6.9 pp, medium confidence
29Kimi K2.674.6%estimated ± 6.9 pp, medium confidence
30Ling 3.0 Flash74.5%measured
31GPT-5.4 nano74.5%estimated ± 6.9 pp, medium confidence
32GPT-5.574.5%estimated ± 6.9 pp, medium confidence
33Muse Spark74.5%estimated ± 6.9 pp, medium confidence
34Granite 4.2 3B74.3%measured
35MiniMax M2.774.3%estimated ± 6.9 pp, medium confidence
36Qwen3.5-122B-A10B74.3%estimated ± 6.9 pp, medium confidence
37Nemotron 3 Nano Omni 30B A3B74.2%measured
38Gemma 4 31B74.2%estimated ± 6.9 pp, medium confidence
39Qwen3.5-27B74.2%estimated ± 6.9 pp, medium confidence
40Ternary Bonsai 2 27B74.0%measured
41GPT-5.274.0%estimated ± 6.9 pp, medium confidence
42GPT-5.3 Codex74.0%estimated ± 6.9 pp, medium confidence
43Ling 3.0 Flash FP873.4%measured
44Nemotron 3.5 Lightning 30B A3B NVFP472.9%measured
45Agents-A172.8%estimated ± 9.7 pp, low confidence
46K-EXAONE 2.072.6%measured
47Command A+72.4%estimated ± 6.9 pp, medium confidence
48GPT-5.472.4%estimated ± 6.9 pp, medium confidence
49Gemma 4 12B72.0%estimated ± 6.9 pp, medium confidence
50GLM-5.271.8%estimated ± 6.9 pp, medium confidence
51GPT-5.4 mini71.8%estimated ± 6.9 pp, medium confidence
52GLM-5-Turbo71.7%estimated ± 6.9 pp, medium confidence
53GPT-5 (high)71.6%estimated ± 6.9 pp, medium confidence
54GPT-5.171.4%estimated ± 6.9 pp, medium confidence
55GPT-5.6 Sol71.2%estimated ± 6.9 pp, medium confidence
56o3-mini71.1%estimated ± 9.7 pp, medium confidence
57Qwen3.5-35B-A3B71.0%estimated ± 6.9 pp, medium confidence
58Gemma 4 26B A4B70.9%estimated ± 6.9 pp, medium confidence
59GLM-570.8%estimated ± 6.9 pp, medium confidence
60Nemotron 3 Super 100B70.0%estimated ± 6.9 pp, medium confidence
61o369.9%estimated ± 6.9 pp, medium confidence
62GPT-5.6 Terra69.7%estimated ± 6.9 pp, medium confidence
63Nemotron 3 Nano 30B69.6%estimated ± 6.9 pp, medium confidence
64GPT-5 (medium)69.2%estimated ± 6.9 pp, medium confidence
65Agents-A1-4B69.1%measured
66Gemini 3 Pro69.0%estimated ± 6.9 pp, medium confidence
67o168.9%estimated ± 6.9 pp, medium confidence
68Kimi K2.568.8%estimated ± 6.9 pp, medium confidence
69Kimi K2.5 (Reasoning)68.8%estimated ± 6.9 pp, medium confidence
70GPT-5.1-Codex68.6%estimated ± 6.9 pp, medium confidence
71GPT-5.1-Codex-Max68.6%estimated ± 6.9 pp, medium confidence
72GPT-OSS 120B67.7%estimated ± 6.9 pp, medium confidence
73MiMo-V2-Pro67.6%estimated ± 6.9 pp, medium confidence
74Mistral Medium 3.5 128B67.6%estimated ± 6.9 pp, medium confidence
75GLM-4.766.8%estimated ± 6.9 pp, medium confidence
76Qwen3.6-27B66.5%estimated ± 6.9 pp, medium confidence
77MiniCPM5-2B66.3%measured
78Step 3.7 Flash66.3%estimated ± 6.9 pp, medium confidence
79GPT-OSS 20B64.6%estimated ± 6.9 pp, medium confidence
80K-Exaone64.3%estimated ± 6.9 pp, medium confidence
81Qwen3.6-35B-A3B64.1%estimated ± 6.9 pp, medium confidence
82Claude Fable 563.5%estimated ± 6.9 pp, medium confidence
83Kimi K2.7 Code63.2%estimated ± 6.9 pp, medium confidence
84Hy3 Preview63.1%measured
85Claude Opus 4.862.6%estimated ± 6.9 pp, medium confidence
86GLM-5V-Turbo62.0%estimated ± 6.9 pp, medium confidence
87Claude Opus 4.7 (Adaptive)60.6%estimated ± 6.9 pp, medium confidence
88Claude Opus 4.5 Thinking60.3%estimated ± 6.9 pp, medium confidence
89North Mini Code60.1%estimated ± 6.9 pp, medium confidence
90Trinity-Large-Preview59.5%estimated ± 6.9 pp, medium confidence
91Trinity-Large-Thinking59.5%estimated ± 6.9 pp, medium confidence
92Claude 4.1 Opus Thinking59.2%estimated ± 6.9 pp, medium confidence
93LFM2.5-2.6B59.2%measured
94Gemini 3 Flash59.1%estimated ± 6.9 pp, medium confidence
95Grok 458.5%estimated ± 6.9 pp, medium confidence
96MiMo-V2-Omni58.5%estimated ± 6.9 pp, medium confidence
97Claude Opus 4.6 (Adaptive)58.3%estimated ± 6.9 pp, medium confidence
98Grok 4.1 Fast (Reasoning)58.2%estimated ± 6.9 pp, medium confidence
99Claude Opus 4.558.0%measured
100Qwen3.5 397B57.9%estimated ± 6.9 pp, medium confidence
101Qwen3.5 397B (Reasoning)57.9%estimated ± 6.9 pp, medium confidence
102Grok 4 Fast (Reasoning)57.6%estimated ± 6.9 pp, medium confidence
103DeepSeek V3.257.2%estimated ± 6.9 pp, medium confidence
104Gemini 2.5 Pro57.1%estimated ± 6.9 pp, medium confidence
105Mistral Small 457.0%estimated ± 6.9 pp, medium confidence
106Mistral Small 4 (Reasoning)57.0%estimated ± 6.9 pp, medium confidence
107Ling 2.6 Flash57.0%measured
108Ultravox v0.6 Llama 3.3 70B56.8%estimated ± 6.9 pp, medium confidence
109Claude 4 Sonnet56.5%estimated ± 6.9 pp, medium confidence
110LFM2.5-8B-A1B56.5%measured
111Claude Opus 4.656.4%estimated ± 6.9 pp, medium confidence
112Gemma 4 E4B56.3%estimated ± 6.9 pp, medium confidence
113Qwen3 Max56.3%estimated ± 6.9 pp, medium confidence
114Claude Opus 4.756.3%estimated ± 6.9 pp, medium confidence
115Qwen3-Omni-30B-A3B-Thinking56.2%estimated ± 6.9 pp, medium confidence
116GPT-4.156.2%estimated ± 6.9 pp, medium confidence
117Llama 4 Maverick56.2%estimated ± 6.9 pp, medium confidence
118DeepSeek V3.1 (Reasoning)56.0%estimated ± 6.9 pp, low confidence
119Kimi K256.0%estimated ± 6.9 pp, low confidence
120Grok Code Fast 156.0%estimated ± 6.9 pp, low confidence
121Claude Sonnet 4.656.0%estimated ± 6.9 pp, low confidence
122DeepSeek V3 032456.0%estimated ± 6.9 pp, low confidence
123MiMo-V2-Flash55.9%estimated ± 6.9 pp, low confidence
124DeepSeek-R155.8%estimated ± 6.9 pp, low confidence
125Llama 4 Scout55.8%estimated ± 6.9 pp, low confidence
126Mistral Medium 355.8%estimated ± 6.9 pp, low confidence
127Gemini 2.5 Flash55.8%estimated ± 6.9 pp, low confidence
128Llama 3.1 405B55.8%estimated ± 6.9 pp, low confidence
129Solar Pro 355.8%measured
130GPT-4.1 mini55.7%estimated ± 6.9 pp, low confidence
131Nemotron Ultra 253B55.7%estimated ± 6.9 pp, low confidence
132Nova Pro55.7%estimated ± 6.9 pp, low confidence
133Gemma 4 E2B55.7%estimated ± 6.9 pp, low confidence
134DeepSeek V3.155.7%estimated ± 6.9 pp, low confidence
135GLM-4.5-Air55.7%estimated ± 6.9 pp, low confidence
136GLM-4.655.6%estimated ± 6.9 pp, low confidence
137Grok 4.1 Fast55.6%estimated ± 6.9 pp, low confidence
138Mistral Large 355.6%estimated ± 6.9 pp, low confidence
139Claude 3 Haiku55.6%estimated ± 6.9 pp, low confidence
140DeepSeek V355.5%estimated ± 6.9 pp, low confidence
141Sarvam 105B55.5%estimated ± 6.9 pp, low confidence
142GPT-4o55.5%estimated ± 6.9 pp, low confidence
143Solar Pro 255.5%estimated ± 6.9 pp, low confidence
144Exaone 4.0 32B55.5%estimated ± 6.9 pp, low confidence
145LFM2.5-VL-1.6B-Extract55.5%estimated ± 6.9 pp, low confidence
146GPT-4.1 nano55.4%estimated ± 6.9 pp, low confidence
147Gemma 3 27B55.4%estimated ± 6.9 pp, low confidence
148Mistral Large 255.4%estimated ± 6.9 pp, low confidence
149Qwen3-Omni-30B-A3B-Instruct55.4%estimated ± 6.9 pp, low confidence
150GPT-4o mini55.4%estimated ± 6.9 pp, low confidence
151Sarvam 30B55.3%estimated ± 6.9 pp, low confidence
152Granite-4.0-H-1B55.3%estimated ± 6.9 pp, low confidence
153Exaone 4.0 1.2B55.3%estimated ± 6.9 pp, low confidence
154Phi-455.3%estimated ± 6.9 pp, low confidence
155DeepSeek R1 Distill Qwen 32B55.3%estimated ± 6.9 pp, low confidence
156Granite-4.0-H-350M55.3%estimated ± 6.9 pp, low confidence
157Granite-4.0-350M55.3%estimated ± 6.9 pp, low confidence
158ZAYA1-8B52.6%measured
159MiniCPM5-1B46.7%measured
160Celeris-144.6%estimated ± 9.7 pp, medium confidence
161LFM2.5-230M38.4%measured
162Mellum2-12B-A2.5B-Thinking35.8%estimated ± 9.7 pp, medium confidence
163Kanana-2 1.3B Instruct34.7%measured
164Mellum2-12B-A2.5B-Instruct34.4%estimated ± 9.7 pp, medium confidence
165Kanana-2 3B Instruct33.3%measured
166LFM2.5-VL-3B25.8%measured
167LLaDA2.2-mini24.9%measured
168LFM2.5-VL-450M12.0%estimated ± 9.7 pp, low confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General