Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

IFBench benchmark: model scores and methodology

An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.

Current published leader
Grok 4.3 (medium)
Top score
83%2-way tie
Primary metric
Prompt-level accuracy · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current IFBench model scores

217 configurations · 162 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 217 current configurations with a reported IFBench score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on IFBench

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. These results sit in a narrow band, so the scale below is zoomed — read each mark against the labelled axis, not the left edge. Marks are placed on the full source precision, so two models sharing a rounded label can still sit at different points. Score labels reproduce Artificial Analysis's rounded display value.

  1. Grok 4.3 medium 83%
  2. MiniMax-M3 83%
  3. Nemotron 3 Ultra 550B A55B Reasoning 81%
  4. Qwen3.7 Max 81%
  5. Nemotron Cascade 2 30B A3B 80%
  6. MiMo-V2.5-Pro 80%
  7. Nova 2.0 Pro Preview low 80%
  8. DeepSeek V4 Flash Reasoning, Max Effort 79%
  9. Qwen3.5 397B A17B Reasoning 79%
  10. Qwen3.7 Plus 78%
  11. Gemini 3.1 Flash-Lite 77%
  12. Gemini 3.1 Pro (Preview) 77%
  13. DeepSeek V4 Pro Reasoning, Max Effort 76%
  14. Gemini 3.5 Flash high 76%
  15. Muse Spark 76%
  16. Qwen3.5 122B A10B Reasoning 76%
  17. Gemma 4 31B Reasoning 76%
  18. GPT-5.3 Codex xhigh 75%
  19. Qwen3.6 Plus 75%
  20. Command A+ 74%
70.7 77.6 84.5

Prompt-level accuracy · higher is better · zoomed scale

Current configurations with a reported IFBench value; missing values are omitted.

How to interpret the result

What do IFBench results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher prompt-level accuracy is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read IFBench as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationPrompt-level accuracy
1 Grok 4.3 (medium)Leader medium 83%
2 MiniMax-M3 Published configuration 83%
3 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 81%
4 Grok 4.3 (low) low 81%
5 Qwen3.7 Max Published configuration 81%
6 Nemotron Cascade 2 30B A3B Published configuration 80%
7 MiMo-V2.5-Pro Published configuration 80%
8 Nova 2.0 Pro Preview (low) low 80%
9 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 79%
10 Nova 2.0 Pro Preview (medium) medium 79%
11 Qwen3.5 397B A17B (Reasoning) Reasoning 79%
12 Qwen3.7 Plus Published configuration 78%
13 Gemini 3.1 Flash-Lite Published configuration 77%
14 Gemini 3.1 Pro (Preview) Published configuration 77%
15 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 76%
16 Gemini 3.5 Flash (high) high 76%
17 Muse Spark Published configuration 76%
18 Qwen3.5 122B A10B (Reasoning) Reasoning 76%
19 Gemma 4 31B (Reasoning) Reasoning 76%
20 GPT-5.3 Codex (xhigh) xhigh 75%
21 Qwen3.6 Plus Published configuration 75%
22 Gemini 3.5 Flash (medium) medium 75%
23 Command A+ Published configuration 74%
24 Gemma 4 12B (Reasoning) Reasoning 74%
25 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 73%
Show the remaining 192 configurations
#ModelConfigurationPrompt-level accuracy
26 GLM-5.2 (max) max 73%
27 GPT-5.6 Sol max 73%
28 Gemma 4 26B A4B (Reasoning) Reasoning 72%
29 MiMo-V2-Flash (Feb 2026) Feb 2026 72%
30 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 71%
31 o3 Published configuration 71%
32 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 71%
33 GPT-5.6 Terra max 71%
34 Solar Pro 3 Published configuration 71%
35 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 71%
36 GPT-5.6 Sol (xhigh) xhigh 71%
37 Nova 2.0 Lite (high) high 71%
38 Mercury 2 Published configuration 70%
39 GPT-5.6 Sol (medium) medium 70%
40 GPT-5.6 Sol (high) high 69%
41 Apriel-v1.6-15B-Thinker Published configuration 69%
42 gpt-oss-120b (high) high 69%
43 Mistral Medium 3.5 Published configuration 69%
44 Nova 2.0 Lite (medium) medium 69%
45 KAT-Coder-Pro V1 Published configuration 68%
46 Qwen3.6 27B (Reasoning) Reasoning 68%
47 MiMo-V2-Omni-0327 Published configuration 67%
48 Step 3.7 Flash Published configuration 67%
49 MiMo-V2.5 Published configuration 67%
50 Qwen3.5 9B (Reasoning) Reasoning 67%
51 KAT Coder Pro V2 Published configuration 67%
52 GPT-5.6 Sol (low) low 67%
53 HyperNova 60B 2605 Published configuration 66%
54 GPT-5.6 Terra (xhigh) xhigh 66%
55 Nex-N2-Pro Published configuration 66%
56 Nova 2.0 Omni (medium) medium 66%
57 Olmo 3.1 32B Think Published configuration 66%
58 gpt-oss-20b (high) high 65%
59 K-EXAONE (Reasoning) Reasoning 65%
60 GPT-5.6 Terra (high) high 64%
61 Qwen3.6 35B A3B (Reasoning) Reasoning 64%
62 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 63%
63 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 63%
64 Hy3-preview (Reasoning) Reasoning 63%
65 Kimi K2.7 Code Published configuration 63%
66 K2 Think V2 Published configuration 63%
67 GPT-5.6 Terra (medium) medium 62%
68 Nova 2.0 Omni (low) low 62%
69 Nova 2.0 Lite (low) low 61%
70 Qwen3 Next 80B A3B (Reasoning) Reasoning 61%
71 K2-V2 (high) high 60%
72 GPT-5.6 Terra (low) low 60%
73 DiffusionGemma 26B A4B Published configuration 59%
74 gpt-oss-120b (low) low 58%
75 NVIDIA Nemotron 3 Nano 4B Published configuration 58%
76 EXAONE 4.5 33B Published configuration 58%
77 gpt-oss-20b (low) low 58%
78 Solar Open 100B (Reasoning) Reasoning 58%
79 North Mini Code Published configuration 58%
80 Ling 2.6 Flash Published configuration 57%
81 Motif-2-12.7B-Reasoning Published configuration 57%
82 Ling-2.6-1T Published configuration 57%
83 Trinity Large Thinking Published configuration 56%
84 K2-V2 (medium) medium 55%
85 Tri-21B-Think Published configuration 55%
86 Falcon-H1R-7B Published configuration 54%
87 Claude 4.5 Haiku (Reasoning) Reasoning 54%
88 MiMo-V2-Omni Published configuration 54%
89 Gemma 4 31B (Non-reasoning) Non-reasoning 53%
90 LFM2.5-8B-A1B Published configuration 53%
91 Jamba Reasoning 3B Published configuration 52%
92 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 52%
93 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 52%
94 Doubao Seed Code Published configuration 51%
95 Qwen3.5 Omni Plus Published configuration 51%
96 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 51%
97 Qwen3.5 4B (Reasoning) Reasoning 50%
98 Step3 VL 10B Published configuration 50%
99 Mi:dm K 2.5 Pro Published configuration 49%
100 MiniCPM5-1B (Reasoning) Reasoning 49%
101 Gemini 2.5 Pro Published configuration 49%
102 Mistral Small 4 (Reasoning) Reasoning 48%
103 Hy3-preview (Non-reasoning) Non-reasoning 48%
104 Grok 4.3 (Non-reasoning) Non-reasoning 48%
105 Gemini 3.5 Flash (minimal) minimal 47%
106 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 47%
107 Llama 3.3 Instruct 70B Published configuration 47%
108 Cogito v2.1 (Reasoning) Reasoning 46%
109 LFM2 24B A2B Published configuration 46%
110 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 46%
111 Qwen3.6 27B (Non-reasoning) Non-reasoning 46%
112 Mi:dm K 2.5 Pro Preview Published configuration 46%
113 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 45%
114 Gemma 4 12B (Non-reasoning) Non-reasoning 45%
115 Ring-2.6-1T Published configuration 45%
116 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 44%
117 Granite 4.1 30B Published configuration 44%
118 Magistral Small 1.2 Published configuration 44%
119 Kimi K2.6 (Non-reasoning) Non-reasoning 44%
120 Qwen3 Omni 30B A3B (Reasoning) Reasoning 43%
121 Ring-flash-2.0 Published configuration 43%
122 LongCat Flash Lite Published configuration 43%
123 Llama 4 Maverick Published configuration 43%
124 Magistral Medium 1.2 Published configuration 43%
125 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 43%
126 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 42%
127 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 42%
128 JT-35B-Flash Published configuration 42%
129 LFM2.5-1.2B-Thinking Published configuration 42%
130 Olmo 3 7B Think Published configuration 41%
131 ERNIE 5.0 Thinking Preview Published configuration 41%
132 Nova 2.0 Omni (Non-reasoning) Non-reasoning 41%
133 LFM2.5-1.2B-Instruct Published configuration 41%
134 K2-V2 (low) low 41%
135 Gemma 4 E4B (Reasoning) Reasoning 41%
136 Gemma 4 E4B (Non-reasoning) Non-reasoning 41%
137 Nova 2.0 Lite (Non-reasoning) Non-reasoning 41%
138 MiMo-V2-Flash (Non-reasoning) Non-reasoning 40%
139 Qwen3 Next 80B A3B Instruct Published configuration 40%
140 K-EXAONE (Non-reasoning) Non-reasoning 40%
141 Llama 4 Scout Published configuration 40%
142 Olmo 3.1 32B Instruct Published configuration 39%
143 ERNIE 4.5 300B A47B Published configuration 39%
144 Llama 3.1 Instruct 405B Published configuration 39%
145 Granite 4.1 8B Published configuration 39%
146 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 38%
147 Devstral 2 Published configuration 38%
148 Qwen3.5 Omni Flash Published configuration 38%
149 HyperCLOVA X SEED Think (32B) 32B 38%
150 Qwen3.5 9B (Non-reasoning) Non-reasoning 38%
151 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 37%
152 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 37%
153 JT-MINI Published configuration 37%
154 Command A Published configuration 36%
155 EXAONE 4.0 32B (Reasoning) Reasoning 36%
156 Mistral Large 3 Published configuration 36%
157 Nova Premier Published configuration 36%
158 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 36%
159 Gemma 4 E2B (Reasoning) Reasoning 36%
160 Nanbeige4.1-3B Published configuration 35%
161 Qwen3 Coder Next Published configuration 35%
162 Jamba 1.7 Large Published configuration 35%
163 MiniCPM5-1B (Non-reasoning) Non-reasoning 35%
164 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 35%
165 Sarvam 105B (high) high 34%
166 INTELLECT-3 Published configuration 34%
167 Granite 4.1 3B Published configuration 34%
168 Gemma 4 E2B (Non-reasoning) Non-reasoning 34%
169 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 33%
170 Qwen3.5 4B (Non-reasoning) Non-reasoning 33%
171 LFM2.5-VL-1.6B Published configuration 33%
172 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 33%
173 Mistral Small 4 (Non-reasoning) Non-reasoning 33%
174 Olmo 3 7B Instruct Published configuration 33%
175 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 33%
176 Ministral 3 14B Published configuration 32%
177 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 32%
178 Granite 4.0 H Small Published configuration 31%
179 Jamba 1.7 Mini Published configuration 31%
180 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 31%
181 Devstral Small 2 Published configuration 31%
182 Qwen3 Omni 30B A3B Instruct Published configuration 31%
183 Llama 3.1 Nemotron Instruct 70B Published configuration 31%
184 Llama 3.2 Instruct 11B (Vision) Vision 30%
185 Qwen3.5 2B (Reasoning) Reasoning 30%
186 Reka Flash 3 Published configuration 30%
187 Nova Micro Published configuration 29%
188 Ministral 3 8B Published configuration 29%
189 Qwen3.5 2B (Non-reasoning) Non-reasoning 29%
190 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 29%
191 Kimi Linear 48B A3B Instruct Published configuration 28%
192 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 28%
193 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 27%
194 Molmo2-8B Published configuration 27%
195 MiniCPM-V 4.6 1.3B Published configuration 27%
196 Sarvam 30B (high) high 26%
197 LFM2 8B A1B Published configuration 26%
198 Apertus 70B Instruct Published configuration 26%
199 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 26%
200 LFM2 2.6B Published configuration 26%
201 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 25%
202 Granite 4.0 H 1B Published configuration 25%
203 Ministral 3 3B Published configuration 24%
204 Ling-mini-2.0 Published configuration 24%
205 Phi-4 Published configuration 24%
206 Exaone 4.0 1.2B (Reasoning) Reasoning 23%
207 Apertus 8B Instruct Published configuration 22%
208 Granite 4.0 Micro Published configuration 22%
209 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 22%
210 Phi-4 Mini Instruct Published configuration 21%
211 Qwen3.5 0.8B (Reasoning) Reasoning 21%
212 Granite 4.0 1B Published configuration 21%
213 Tiny Aya Global Published configuration 20%
214 Molmo 7B-D Published configuration 20%
215 Granite 4.0 H 350M Published configuration 17%
216 Granite 4.0 350M Published configuration 15%
217 Gemma 3 270M Published configuration 12%

Showing the top 25 of 217 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

An instruction-following benchmark with diverse, verifiable out-of-domain output constraints.

How to read the score

The published metric is Prompt-level accuracy. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same ifbench field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official IFBench resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗