Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of precise, checkable formatting instructions followed, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
288 configurations of 195 models
Unit
% of instructions
Licence
Artificial Analysis commercial data licence

IFBench: Qwen3.5-2B

Top 15 of 195 results · % of instructions, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Grok 4.3 (medium reasoning)xAI 83.3%
  2. 2 Grok 4.20 (0309, reasoning)xAI 82.9%
  3. 2 MiniMax-M3MiniMax 82.9%
  4. 4 Grok 4.20xAI 81.2%
  5. 5 Qwen3.7-MaxAlibaba 80.5%
  6. 6 MiMo-V2.5-Pro (reasoning on)Xiaomi 79.9%
  7. 7 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek 79.2%
  8. 8 Qwen3.5-397B-A17BAlibaba 78.8%
  9. 9 Qwen3.7-PlusAlibaba 78.0%
  10. 9 Gemini 3 Flash Preview (reasoning on)Google 78.0%
  11. 11 GPT-5.2-Codex (extra-high reasoning)OpenAI 77.6%
  12. 12 Gemini 3.1 Flash-Lite PreviewGoogle 77.2%
  13. 13 Gemini 3.1 Pro PreviewGoogle 77.1%
  14. 14 Qwen3.6-Max-PreviewAlibaba 76.6%
  15. 15 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek 76.5%
  16. 176 Qwen3.5-2B (reasoning on)Alibaba 31.5%

Full results

Artificial Analysis: IFBench, % of instructions, higher is better
#ModelIFBench
% of instructions, higher is better
Price
$ per million tokens, in / out
1 Grok 4.3 (medium reasoning)xAI · best of 4 settings
83.3%
$1.25 / $2.50
2 Grok 4.20 (0309, reasoning)xAI
82.9%
–
2 MiniMax-M3MiniMax
82.9%
$0.30 / $1.20
4 Grok 4.20xAI · best of 2 settings
81.2%
$1.25 / $2.50
5 Qwen3.7-MaxAlibaba
80.5%
$1.48 / $4.42
6 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
79.9%
$0.43 / $0.87
7 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek · best of 3 settings
79.2%
$0.14 / $0.28
8 Qwen3.5-397B-A17BAlibaba · best of 2 settings
78.8%
$0.55 / $3.50
9 Qwen3.7-PlusAlibaba
78.0%
$0.32 / $1.28
9 Gemini 3 Flash Preview (reasoning on)Google · best of 2 settings
78.0%
$0.50 / $3
11 GPT-5.2-Codex (extra-high reasoning)OpenAI
77.6%
$1.75 / $14
12 Gemini 3.1 Flash-Lite PreviewGoogle
77.2%
$0.25 / $1.50
13 Gemini 3.1 Pro PreviewGoogle
77.1%
$2 / $12
14 Qwen3.6-Max-PreviewAlibaba
76.6%
$1.03 / $6.16
15 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek · best of 3 settings
76.5%
$1.42 / $2.83
16 Gemini 3.5 Flash (high reasoning)Google · best of 3 settings
76.3%
$1.50 / $9
16 GLM-5.1Z.ai · best of 2 settings
76.3%
$1.38 / $4.40
18 Kimi K2.6Moonshot AI · best of 2 settings
76.0%
$0.95 / $4
19 Muse SparkMeta
75.9%
–
19 GPT-5.4 nano (extra-high reasoning)OpenAI · best of 3 settings
75.9%
$0.20 / $1.25
19 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
75.9%
$5 / $30
22 Qwen3.5-122B-A10BAlibaba · best of 2 settings
75.7%
$0.26 / $2.08
22 MiniMax-M2.7MiniMax
75.7%
$0.30 / $1.20
24 Qwen3.5-27BAlibaba · best of 2 settings
75.6%
$0.27 / $2.16
24 Gemma 4 31BGoogle · best of 2 settings
75.6%
$0.14 / $0.40
26 GPT-5.2 (extra-high reasoning)OpenAI · best of 3 settings
75.4%
$1.75 / $14
26 GPT-5 mini (high reasoning)OpenAI · best of 3 settings
75.4%
$0.25 / $2
26 GPT-5.3-Codex (extra-high reasoning)OpenAI
75.4%
$1.75 / $14
29 Qwen3.6-PlusAlibaba
75.2%
$0.33 / $1.95
30 GPT-5-Codex (high reasoning)OpenAI
74.2%
–
31 GPT-5.4 (extra-high reasoning)OpenAI · best of 3 settings
73.9%
$2.50 / $15
32 Gemma 4 12BGoogle · best of 2 settings
73.5%
–
33 GLM-5.2 (max reasoning)Z.ai
73.3%
$1.40 / $4.40
33 GPT-5.4 mini (extra-high reasoning)OpenAI · best of 3 settings
73.3%
$0.75 / $4.50
35 GLM-5-TurboZ.ai
73.2%
$1.20 / $4
36 GPT-5 (high reasoning)OpenAI · best of 4 settings
73.1%
$1.25 / $10
37 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
72.9%
$1.25 / $10
38 GPT-5.6 Sol (max reasoning)OpenAI · best of 5 settings
72.7%
$4 / $20
39 Qwen3.5-35B-A3BAlibaba · best of 2 settings
72.5%
$0.16 / $1.30
40 Gemma 4 26B A4BGoogle · best of 2 settings
72.4%
$0.10 / $0.30
41 MiniMax-M2MiniMax
72.3%
$0.30 / $1.20
41 GLM-5Z.ai · best of 2 settings
72.3%
$0.95 / $2.55
43 MiniMax-M2.5MiniMax
71.6%
$0.30 / $1.20
44 Nemotron 3 Super 120B A12BNVIDIA
71.5%
$0.085 / $0.40
44 GPT-5.5 Instant (2026-05-26)OpenAI
71.5%
–
46 o3OpenAI
71.4%
$2 / $8
47 GPT-5.6 Terra (max reasoning)OpenAI · best of 5 settings
71.2%
$2 / $12
47 Solar Pro 3Upstage
71.2%
$0.15 / $0.60
49 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA · best of 2 settings
71.1%
$0.05 / $0.20
50 Qwen3-Max-ThinkingAlibaba
70.7%
$0.78 / $3.90
50 Nova 2 Lite (high reasoning)Amazon · best of 4 settings
70.7%
$0.30 / $2.50
52 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
70.4%
–
53 o1OpenAI
70.3%
$15 / $60
54 Kimi K2.5Moonshot AI · best of 2 settings
70.2%
$0.57 / $2.85
55 GPT-5.1-Codex (high reasoning)OpenAI
70.0%
$1.25 / $10
56 MiniMax-M2.1MiniMax
69.9%
$0.30 / $1.20
57 Mercury 2Inception
69.8%
$0.25 / $0.75
58 gpt-oss-120b (high reasoning)OpenAI · best of 2 settings
69.0%
$0.15 / $0.60
59 Mistral Medium 3.5Mistral AI
68.8%
$1.50 / $7.50
60 o4-mini (high reasoning)OpenAI
68.7%
$1.10 / $4.40
61 Kimi K2 ThinkingMoonshot AI
68.1%
$0.60 / $2.50
62 GPT-5.1-Codex-Mini (high reasoning)OpenAI
67.9%
$0.25 / $2
62 GLM-4.7Z.ai · best of 2 settings
67.9%
$0.54 / $1.98
64 Qwen3.6-27BAlibaba · best of 2 settings
67.6%
$0.30 / $3.20
64 GPT-5 nano (high reasoning)OpenAI · best of 3 settings
67.6%
$0.05 / $0.40
66 o3-mini (high reasoning)OpenAI
67.1%
$1.10 / $4.40
66 MiMo-V2.5Xiaomi
67.1%
$0.17 / $0.34
68 Qwen3.5-9BAlibaba · best of 2 settings
66.7%
$0.10 / $0.15
69 gpt-oss-20b (high reasoning)OpenAI · best of 2 settings
65.1%
$0.03 / $0.15
70 Step 3.5 FlashStepFun
64.6%
$0.10 / $0.30
71 Qwen3.6-35B-A3BAlibaba · best of 2 settings
64.4%
$0.10 / $1
72 DeepSeek-V3.2-SpecialeDeepSeek
63.9%
–
73 Claude Fable 5 (max reasoning)Anthropic
63.5%
$10 / $50
74 Kimi K2.7 CodeMoonshot AI
63.1%
$0.95 / $4
74 Hy3 Preview (reasoning on)Tencent · best of 2 settings
63.1%
$0.18 / $0.60
76 Claude Opus 4.8 (max reasoning)Anthropic
62.2%
$5 / $25
77 GLM-5V-TurboZ.ai
61.1%
$1.20 / $4
78 GLM-4.7-FlashZ.ai · best of 2 settings
60.8%
$0.06 / $0.40
79 Qwen3-Next-80B-A3B-ThinkingAlibaba
60.7%
$0.15 / $1.20
79 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
60.7%
$0.30 / $0.96
81 DiffusionGemma 26B A4BGoogle
59.5%
–
82 Qwen3-VL-32B-ThinkingAlibaba
59.4%
–
83 Claude Opus 4.7 (max reasoning)Anthropic · best of 2 settings
58.6%
$5 / $25
84 Claude Opus 4.5 (reasoning on)Anthropic · best of 2 settings
58.0%
$5 / $25
85 Claude Sonnet 4.5 (reasoning on)Anthropic · best of 2 settings
57.3%
$3 / $15
86 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek · best of 2 settings
57.0%
$0.27 / $1
87 Claude Sonnet 4.6 (max reasoning)Anthropic · best of 2 settings
56.6%
$3 / $15
88 Qwen3-VL-235B-A22B-ThinkingAlibaba
56.5%
$0.40 / $4
89 Trinity Large ThinkingArcee AI
56.3%
$0.25 / $0.80
90 Claude Opus 4.1 (reasoning on)Anthropic
55.4%
$15 / $75
91 Claude Sonnet 4 (reasoning on)Anthropic · best of 2 settings
54.7%
$3 / $15
92 Claude Haiku 4.5 (reasoning on)Anthropic · best of 2 settings
54.3%
$1 / $5
93 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
54.2%
$0.27 / $0.41
94 Qwen3-Max-Thinking-PreviewAlibaba
53.8%
–
95 Grok 4xAI
53.7%
–
96 Claude Opus 4.6 (max reasoning)Anthropic · best of 2 settings
53.1%
$5 / $25
97 Grok 4.1 Fast (reasoning)xAI
52.7%
–
98 Gemini 2.5 Flash-Lite Preview (09-2025, reasoning on)Google · best of 2 settings
52.6%
–
99 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google · best of 2 settings
52.3%
–
100 Qwen3.5-4B (reasoning on)Alibaba · best of 2 settings
52.0%
–
101 Qwen3-235B-A22B-Thinking-2507Alibaba
51.2%
$0.30 / $3
101 Qwen3.5-Omni-PlusAlibaba
51.2%
–
103 Qwen3-30B-A3B-Thinking-2507Alibaba
50.7%
$0.20 / $2.40
104 Grok 4 Fast (reasoning)xAI
50.5%
–
105 Gemini 2.5 Flash (reasoning on)Google · best of 2 settings
50.3%
$0.30 / $2.50
106 Gemini 2.5 Flash-Lite (reasoning on)Google · best of 2 settings
49.9%
$0.10 / $0.40
107 Qwen3-4B-Thinking-2507Alibaba
49.8%
–
108 Gemini 2.5 ProGoogle
48.7%
$1.25 / $10
109 Mistral Small 4Mistral AI · best of 2 settings
48.2%
$0.15 / $0.60
110 Qwen3-Max-PreviewAlibaba
48.0%
–
111 Grok 4.20 (0309, non-reasoning, no reasoning)xAI
47.8%
–
112 Llama 3.3 70B InstructMeta
47.1%
$0.59 / $0.79
113 Qwen3-235B-A22B-Instruct-2507Alibaba
46.1%
$0.15 / $0.75
114 Qwen3-VL-30B-A3B-ThinkingAlibaba
45.1%
$0.29 / $1
115 GPT-5 ChatOpenAI
45.0%
–
116 Magistral Small 1.2Mistral AI
44.4%
–
117 Gemma 4 E4BGoogle · best of 2 settings
44.2%
–
117 Qwen3-MaxAlibaba
44.2%
$0.78 / $3.90
119 GLM-4.5Z.ai
44.1%
$0.60 / $2.20
120 Qwen3-Omni-30B-A3B-ThinkingAlibaba
43.4%
–
120 GLM-4.6 (reasoning on)Z.ai · best of 2 settings
43.4%
$0.50 / $2
122 Llama 4 MaverickMeta
43.0%
$0.27 / $0.85
122 Magistral Medium 1.2Mistral AI
43.0%
–
122 GPT-4.1OpenAI
43.0%
$2 / $8
125 Qwen3-VL-235B-A22B-InstructAlibaba
42.7%
$0.30 / $1.50
126 Kimi K2 (0905)Moonshot AI
41.7%
$0.60 / $2.50
127 Qwen3-30B-A3B (reasoning on)Alibaba · best of 2 settings
41.5%
$0.12 / $0.50
127 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
41.5%
$0.55 / $1.65
127 Kimi K2 (0711)Moonshot AI
41.5%
$0.57 / $2.30
130 Grok Code Fast 1xAI
41.4%
–
131 DeepSeek-V3-0324DeepSeek
41.0%
$0.25 / $1
132 Qwen3-14B (reasoning on)Alibaba · best of 2 settings
40.5%
$0.12 / $0.24
132 Qwen3-Coder-480B-A35BAlibaba
40.5%
$0.35 / $1.50
134 Qwen3-VL-8B-ThinkingAlibaba
39.9%
$0.18 / $2.10
135 Mistral Medium 3.1Mistral AI
39.8%
$0.40 / $2
136 Qwen3-Next-80B-A3B-InstructAlibaba
39.7%
$0.10 / $1.10
137 DeepSeek-R1-0528DeepSeek
39.6%
$0.50 / $2.18
138 Llama 4 ScoutMeta
39.5%
$0.18 / $0.59
139 Mistral Medium 3Mistral AI
39.3%
$0.40 / $2
140 Qwen3-VL-32B-InstructAlibaba
39.2%
$0.10 / $0.42
141 DeepSeek-R1DeepSeek
39.0%
$0.70 / $2.50
142 Qwen3-235B-A22B (reasoning on)Alibaba · best of 2 settings
38.7%
$0.46 / $1.82
143 GPT-4.1 miniOpenAI
38.3%
$0.40 / $1.60
144 Nova Pro 1.0Amazon
38.1%
$0.80 / $3.20
144 Devstral 2Mistral AI
38.1%
$0.40 / $2
146 Gemma 4 E2BGoogle · best of 2 settings
38.0%
–
146 Qwen3.5-Omni-FlashAlibaba
38.0%
–
148 Grok 4 Fast (non-reasoning, no reasoning)xAI
37.7%
–
149 GLM-4.5-AirZ.ai
37.6%
$0.14 / $0.86
150 Qwen2.5-72B-InstructAlibaba
36.9%
$0.36 / $0.40
151 Gemma 3 12BGoogle
36.7%
$0.05 / $0.15
152 Qwen3-VL-4B-ThinkingAlibaba
36.6%
–
153 Command ACohere
36.5%
$2.50 / $10
153 Grok 4.1 Fast (non-reasoning, no reasoning)xAI
36.5%
–
155 Qwen3-32B (reasoning on)Alibaba · best of 2 settings
36.3%
$0.14 / $0.40
156 Mistral Large 3Mistral AI
36.2%
$0.50 / $1.50
157 GPT-4o (2024-08-06)OpenAI
36.0%
$2.50 / $10
158 Qwen3-Coder-NextAlibaba
35.2%
$0.18 / $0.90
159 DeepSeek-V3DeepSeek
34.8%
$0.26 / $1.03
160 Devstral Small 1.1Mistral AI
34.6%
–
161 Llama 3.1 70B InstructMeta
34.4%
$0.40 / $0.40
162 GLM-4.5V (reasoning on)Z.ai · best of 2 settings
34.2%
$0.60 / $1.80
162 Nova Lite 1.0Amazon
34.2%
$0.06 / $0.24
164 Qwen3-4B-Instruct-2507Alibaba
33.5%
–
164 Qwen3-8B (reasoning on)Alibaba · best of 2 settings
33.5%
$0.12 / $0.46
164 Mistral Small 3.2 24BMistral AI
33.5%
$0.094 / $0.25
167 Qwen3-VL-30B-A3B-InstructAlibaba
33.1%
$0.15 / $0.60
167 Qwen3-30B-A3B-Instruct-2507Alibaba
33.1%
$0.09 / $0.30
169 Qwen3-Coder-30B-A3B-InstructAlibaba
32.7%
$0.07 / $0.28
170 Qwen3-VL-8B-InstructAlibaba
32.3%
$0.12 / $0.46
171 Ministral 3 14BMistral AI
32.0%
$0.20 / $0.20
171 GPT-4.1 nanoOpenAI
32.0%
$0.10 / $0.40
173 Qwen3-VL-4B-InstructAlibaba
31.8%
–
173 Gemma 3 27BGoogle
31.8%
$0.12 / $0.20
175 Mistral Large 2 (2407)Mistral AI
31.6%
$2 / $6
176 Qwen3.5-2B (reasoning on)Alibaba · best of 2 settings
31.5%
–
177 Qwen3-Omni-30B-A3B-InstructAlibaba
31.2%
–
177 Devstral Small 2Mistral AI
31.2%
–
179 GPT-4o mini (2024-07-18)OpenAI
31.0%
$0.15 / $0.60
180 Reka Flash 3Reka AI
30.4%
$0.10 / $0.20
181 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
30.1%
$0.30 / $0.90
182 Devstral MediumMistral AI
29.9%
–
182 Mistral Small 3.1 24BMistral AI
29.9%
$0.35 / $0.56
184 Nova Micro 1.0Amazon
29.4%
$0.035 / $0.14
185 Ministral 3 8BMistral AI
29.1%
$0.15 / $0.15
186 Llama 3.1 8B InstructMeta
28.6%
$0.05 / $0.08
187 Gemma 3 4BGoogle
28.3%
$0.05 / $0.10
188 Kimi Linear 48B A3B InstructMoonshot AI
28.1%
–
189 Ministral 3 3BMistral AI
26.8%
$0.10 / $0.10
190 Mistral Small 3Mistral AI
26.4%
$0.05 / $0.08
191 Llama 3.2 3B InstructMeta
26.2%
$0.05 / $0.33
192 Phi-4Microsoft
23.5%
$0.07 / $0.14
193 Llama 3.2 1B InstructMeta
22.8%
$0.027 / $0.20
194 Qwen3.5-0.8B (no reasoning)Alibaba · best of 2 settings
21.6%
–
195 Gemma 3 270MGoogle
12.1%
–

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of precise, checkable formatting instructions followed, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not judgement about what an instruction meant.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.