benchgap
Instruction following

IFEval leaderboard

As of 2026-10-07, the highest measured score on IFEval is 95.0% by Qwen3.5-27B. 135 more models have estimated scores, calibrated from the benchmarks they were measured on.

Measured scores: benchlm.ai.

#ModelScoreSource
1Qwen3.5-27B95.0%measured
2Agents-A194.8%measured
3Agents-A1-4B94.8%measured
4Qwen3.7 Plus94.6%measured
5Qwen3.6 Plus94.3%measured
6Qwen3.7 Max94.3%measured
7MiniMax M393.9%estimated ± 1.3 pp, medium confidence
8dots3-note Preview93.9%measured
9Kimi K2.593.9%measured
10o3-mini93.9%measured
11Nemotron 3 Ultra93.9%estimated ± 1.3 pp, medium confidence
12Grok 4.393.9%estimated ± 1.3 pp, medium confidence
13MiMo-V2.5-Pro93.8%estimated ± 1.3 pp, high confidence
14GPT-5.2-Codex93.7%estimated ± 1.3 pp, high confidence
15Gemini 3.1 Pro93.7%estimated ± 1.3 pp, high confidence
16Qwen 3.6 Max (preview)93.7%estimated ± 1.3 pp, high confidence
17DeepSeek V4 Pro 081393.7%estimated ± 1.3 pp, high confidence
18Gemini 3.5 Flash93.7%estimated ± 1.3 pp, high confidence
19GLM-5.193.7%estimated ± 1.3 pp, high confidence
20Kimi K2.693.7%estimated ± 1.3 pp, high confidence
21GPT-5.4 nano93.7%estimated ± 1.3 pp, high confidence
22GPT-5.593.7%estimated ± 1.3 pp, high confidence
23Muse Spark93.7%estimated ± 1.3 pp, high confidence
24MiniMax M2.793.7%estimated ± 1.3 pp, high confidence
25Gemma 4 31B93.7%estimated ± 1.3 pp, high confidence
26GPT-5.293.7%estimated ± 1.3 pp, high confidence
27GPT-5.3 Codex93.7%estimated ± 1.3 pp, high confidence
28Command A+93.6%estimated ± 1.3 pp, high confidence
29GPT-5.493.6%estimated ± 1.3 pp, high confidence
30Gemma 4 12B93.6%estimated ± 1.3 pp, high confidence
31GLM-5.293.6%estimated ± 1.3 pp, high confidence
32GPT-5.4 mini93.6%estimated ± 1.3 pp, high confidence
33GLM-5-Turbo93.6%estimated ± 1.3 pp, high confidence
34GPT-5 (high)93.6%estimated ± 1.3 pp, high confidence
35GPT-5.193.6%estimated ± 1.3 pp, high confidence
36GPT-5.6 Sol93.5%estimated ± 1.3 pp, high confidence
37Gemma 4 26B A4B93.5%estimated ± 1.3 pp, high confidence
38Nemotron 3 Super 100B93.5%estimated ± 1.3 pp, high confidence
39o393.5%estimated ± 1.3 pp, high confidence
40GPT-5.6 Terra93.5%estimated ± 1.3 pp, high confidence
41Solar Pro 393.5%estimated ± 1.3 pp, high confidence
42Nemotron 3 Nano 30B93.5%estimated ± 1.3 pp, high confidence
43GPT-5 (medium)93.4%estimated ± 1.3 pp, high confidence
44Gemini 3 Pro93.4%estimated ± 1.3 pp, high confidence
45Kimi K2.5 (Reasoning)93.4%estimated ± 1.3 pp, high confidence
46GPT-5.1-Codex93.4%estimated ± 1.3 pp, high confidence
47GPT-5.1-Codex-Max93.4%estimated ± 1.3 pp, high confidence
48Qwen3.5-122B-A10B93.4%measured
49GPT-OSS 120B93.4%estimated ± 1.3 pp, high confidence
50MiMo-V2-Pro93.3%estimated ± 1.3 pp, high confidence
51Mistral Medium 3.5 128B93.3%estimated ± 1.3 pp, high confidence
52GLM-4.793.3%estimated ± 1.3 pp, high confidence
53Qwen3.6-27B93.3%estimated ± 1.3 pp, high confidence
54Step 3.7 Flash93.3%estimated ± 1.3 pp, high confidence
55GPT-OSS 20B93.1%estimated ± 1.3 pp, high confidence
56K-Exaone93.1%estimated ± 1.3 pp, high confidence
57Qwen3.6-35B-A3B93.1%estimated ± 1.3 pp, high confidence
58Claude Fable 593.0%estimated ± 1.3 pp, high confidence
59Nemotron 3 Nano Omni 30B A3B93.0%estimated ± 1.3 pp, high confidence
60Kimi K2.7 Code93.0%estimated ± 1.3 pp, high confidence
61Claude Opus 4.892.9%estimated ± 1.3 pp, high confidence
62Inkling-Small92.8%estimated ± 3.5 pp, medium confidence
63MAI-Thinking-192.8%estimated ± 3.5 pp, medium confidence
64Qwen3.8-Flash-Next92.8%estimated ± 3.5 pp, medium confidence
65Qwen3.8 Max92.8%estimated ± 3.5 pp, medium confidence
66Qwen3.8-Omni-Flash92.8%estimated ± 3.5 pp, medium confidence
67Inkling92.8%estimated ± 3.5 pp, high confidence
68Qwen3.8-27B92.8%estimated ± 3.5 pp, high confidence
69Beam92.8%estimated ± 3.5 pp, high confidence
70Solar Open 292.8%estimated ± 3.5 pp, high confidence
71Granite 4.2 8B92.8%estimated ± 3.5 pp, high confidence
72Granite 4.2 30B92.8%estimated ± 3.5 pp, high confidence
73Mercury 2.592.8%estimated ± 3.5 pp, high confidence
74Muse Glimmer 30B92.8%estimated ± 3.5 pp, high confidence
75A.X K292.8%estimated ± 3.5 pp, high confidence
76Ling 3.0 Flash92.8%estimated ± 3.5 pp, high confidence
77Granite 4.2 3B92.8%estimated ± 3.5 pp, high confidence
78Ling 3.0 Flash FP892.8%estimated ± 3.5 pp, high confidence
79Nemotron 3.5 Lightning 30B A3B NVFP492.8%estimated ± 3.5 pp, high confidence
80GLM-5V-Turbo92.8%estimated ± 1.3 pp, high confidence
81Hy3 Preview92.7%estimated ± 3.5 pp, high confidence
82GLM-592.6%measured
83Qwen3.5 397B92.6%measured
84Claude Opus 4.7 (Adaptive)92.6%estimated ± 1.3 pp, high confidence
85Claude Opus 4.5 Thinking92.5%estimated ± 1.3 pp, high confidence
86North Mini Code92.5%estimated ± 1.3 pp, high confidence
87Ling 2.6 Flash92.4%estimated ± 1.3 pp, high confidence
88K-EXAONE 2.092.4%measured
89Trinity-Large-Preview92.3%estimated ± 1.3 pp, high confidence
90Trinity-Large-Thinking92.3%estimated ± 1.3 pp, high confidence
91o192.2%measured
92Claude 4.1 Opus Thinking92.2%estimated ± 1.3 pp, high confidence
93Gemini 3 Flash92.2%estimated ± 1.3 pp, high confidence
94LFM2.5-2.6B92.0%estimated ± 3.5 pp, high confidence
95Grok 492.0%estimated ± 1.3 pp, high confidence
96MiMo-V2-Omni91.9%estimated ± 1.3 pp, high confidence
97Qwen3.5-35B-A3B91.9%measured
98Claude Opus 4.6 (Adaptive)91.9%estimated ± 1.3 pp, high confidence
99LFM2.5-8B-A1B91.8%measured
100Grok 4.1 Fast (Reasoning)91.8%estimated ± 1.3 pp, high confidence
101Qwen3.5 397B (Reasoning)91.6%estimated ± 1.3 pp, high confidence
102Grok 4 Fast (Reasoning)91.4%estimated ± 1.3 pp, high confidence
103Ternary Bonsai 2 27B91.3%measured
104DeepSeek V3.291.2%estimated ± 1.3 pp, high confidence
105Gemini 2.5 Pro91.1%estimated ± 1.3 pp, high confidence
106Mistral Small 491.0%estimated ± 1.3 pp, high confidence
107Mistral Small 4 (Reasoning)91.0%estimated ± 1.3 pp, high confidence
108Claude Opus 4.590.9%measured
109Ultravox v0.6 Llama 3.3 70B90.7%estimated ± 1.3 pp, high confidence
110Claude 4 Sonnet90.3%estimated ± 1.3 pp, high confidence
111Claude Opus 4.690.1%estimated ± 1.3 pp, high confidence
112Gemma 4 E4B90.0%estimated ± 1.3 pp, high confidence
113Qwen3 Max90.0%estimated ± 1.3 pp, high confidence
114Claude Opus 4.789.8%estimated ± 1.3 pp, high confidence
115Qwen3-Omni-30B-A3B-Thinking89.8%estimated ± 1.3 pp, high confidence
116Llama 4 Maverick89.6%estimated ± 1.3 pp, high confidence
117DeepSeek V3.1 (Reasoning)89.1%estimated ± 1.3 pp, high confidence
118Kimi K289.1%estimated ± 1.3 pp, high confidence
119Grok Code Fast 189.1%estimated ± 1.3 pp, high confidence
120Claude Sonnet 4.689.0%estimated ± 1.3 pp, high confidence
121DeepSeek V3 032488.9%estimated ± 1.3 pp, high confidence
122GPT-4.1 mini88.5%measured
123MiMo-V2-Flash88.5%estimated ± 1.3 pp, high confidence
124DeepSeek-R188.4%estimated ± 1.3 pp, high confidence
125Llama 4 Scout88.3%estimated ± 1.3 pp, high confidence
126Mistral Medium 388.2%estimated ± 1.3 pp, high confidence
127Gemini 2.5 Flash88.1%estimated ± 1.3 pp, high confidence
128Llama 3.1 405B88.1%estimated ± 1.3 pp, high confidence
129Nemotron Ultra 253B87.7%estimated ± 1.3 pp, high confidence
130Nova Pro87.7%estimated ± 1.3 pp, high confidence
131Gemma 4 E2B87.6%estimated ± 1.3 pp, high confidence
132DeepSeek V3.187.5%estimated ± 1.3 pp, high confidence
133GLM-4.5-Air87.4%estimated ± 1.3 pp, high confidence
134GPT-4.187.4%measured
135GLM-4.686.9%estimated ± 1.3 pp, high confidence
136Grok 4.1 Fast86.8%estimated ± 1.3 pp, high confidence
137MiniCPM5-2B86.7%measured
138Mistral Large 386.6%estimated ± 1.3 pp, high confidence
139Claude 3 Haiku86.6%estimated ± 1.3 pp, high confidence
140DeepSeek V386.1%measured
141ZAYA1-8B85.6%measured
142Sarvam 105B85.5%estimated ± 1.3 pp, high confidence
143GPT-4o85.4%estimated ± 1.3 pp, high confidence
144Solar Pro 285.0%estimated ± 1.3 pp, high confidence
145Exaone 4.0 32B84.8%estimated ± 1.3 pp, high confidence
146LFM2.5-VL-1.6B-Extract84.5%estimated ± 1.3 pp, high confidence
147Gemma 3 27B83.4%estimated ± 1.3 pp, medium confidence
148GPT-4.1 nano83.2%measured
149Mistral Large 282.9%estimated ± 1.3 pp, medium confidence
150Qwen3-Omni-30B-A3B-Instruct82.9%estimated ± 1.3 pp, medium confidence
151GPT-4o mini82.7%estimated ± 1.3 pp, medium confidence
152LFM2.5-VL-3B82.3%measured
153Kanana-2 3B Instruct81.0%measured
154Celeris-180.8%measured
155MiniCPM5-1B80.4%measured
156LLaDA2.2-mini78.3%estimated ± 3.5 pp, medium confidence
157Kanana-2 1.3B Instruct77.6%measured
158Sarvam 30B77.1%estimated ± 1.3 pp, medium confidence
159Granite-4.0-H-1B76.6%estimated ± 1.3 pp, medium confidence
160Mellum2-12B-A2.5B-Thinking76.5%measured
161Mellum2-12B-A2.5B-Instruct75.8%measured
162Exaone 4.0 1.2B75.1%estimated ± 1.3 pp, medium confidence
163LFM2.5-230M71.7%measured
164Phi-471.6%estimated ± 1.3 pp, medium confidence
165DeepSeek R1 Distill Qwen 32B70.3%estimated ± 1.3 pp, medium confidence
166LFM2.5-VL-450M61.2%measured
167Granite-4.0-H-350M54.4%estimated ± 1.3 pp, medium confidence
168Granite-4.0-350M47.5%estimated ± 1.3 pp, medium confidence

All leaderboards

Agentic · terminal

Agentic · tools

Coding

Math

Knowledge & reasoning

Instruction following

Multilingual

Vision & documents

Long context

General