Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

AA-Omniscience Non-Hallucination Rate benchmark: model scores and methodology

One minus the AA-Omniscience hallucination rate, where incorrect answers are divided by incorrect, partial, and not-attempted outcomes.

Current published leader
MiniCPM5-1B (Non-reasoning)
Top score
99%
Primary metric
Non-hallucination rate · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current AA-Omniscience Non-Hallucination Rate model scores

244 configurations · 176 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 244 current configurations with a reported AA-Omniscience Non-Hallucination Rate score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on AA-Omniscience Non-Hallucination Rate

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. MiniCPM5-1B Non-reasoning 99%
  2. G9v3-3B 88%
  3. Command A+ 86%
  4. Grok 4.3 medium 84%
  5. MiniMax-M3 84%
  6. Qwen3.7 Max 77%
  7. MiMo-V2.5-Pro 75%
  8. Claude 4.5 Haiku Non-reasoning 75%
  9. Qwen3.7 Plus 75%
  10. GLM-5.2 max 72%
  11. Nemotron 3 Ultra 550B A55B Reasoning 71%
  12. Gemma 4 E4B Reasoning 70%
  13. Qwen3.6 Plus 68%
  14. MiMo-V2.5 68%
  15. Gemma 3 270M 68%
  16. Gemma 4 E2B Reasoning 67%
  17. Gemini 3.5 Flash-Lite 66%
  18. Qwen3.5 Omni Plus 64%
  19. Claude Sonnet 5 Adaptive Reasoning, Max Effort 63%
  20. Muse Spark 1.1 xhigh 62%
0.0 49.6 99.1

Non-hallucination rate · higher is better

Current configurations with a reported AA-Omniscience Non-Hallucination Rate value; missing values are omitted.

How to interpret the result

What do AA-Omniscience Non-Hallucination Rate results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher non-hallucination rate is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read AA-Omniscience Non-Hallucination Rate as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationNon-hallucination rate
1 MiniCPM5-1B (Non-reasoning)Leader Non-reasoning 99%
2 G9v3-3B Published configuration 88%
3 Command A+ Published configuration 86%
4 Grok 4.3 (medium) medium 84%
5 MiniMax-M3 Published configuration 84%
6 Grok 4.3 (low) low 84%
7 MiniCPM5-1B (Reasoning) Reasoning 81%
8 Qwen3.7 Max Published configuration 77%
9 MiMo-V2.5-Pro Published configuration 75%
10 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 75%
11 Qwen3.7 Plus Published configuration 75%
12 Claude 4.5 Haiku (Reasoning) Reasoning 74%
13 GLM-5.2 (max) max 72%
14 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 71%
15 Gemma 4 E4B (Reasoning) Reasoning 70%
16 Qwen3.6 Plus Published configuration 68%
17 MiMo-V2.5 Published configuration 68%
18 Gemma 3 270M Published configuration 68%
19 Gemma 4 E2B (Reasoning) Reasoning 67%
20 Gemini 3.5 Flash-Lite Published configuration 66%
21 GLM-5.2 (Non-reasoning) Non-reasoning 66%
22 Qwen3.5 Omni Plus Published configuration 64%
23 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 63%
24 Muse Spark 1.1 (xhigh) xhigh 62%
25 MiMo-V2-Omni-0327 Published configuration 60%
Show the remaining 219 configurations
#ModelConfigurationNon-hallucination rate
26 Qwen3.5 0.8B (Reasoning) Reasoning 60%
27 Kimi K2.6 (Non-reasoning) Non-reasoning 57%
28 JT-4.1 Flash 236B A21B Published configuration 57%
29 MiMo-V2-Omni Published configuration 56%
30 Qwen3.5 2B (Reasoning) Reasoning 53%
31 Qwen3.6 27B (Reasoning) Reasoning 52%
32 MiMo-V2-Flash (Feb 2026) Feb 2026 52%
33 Motif 3 (Beta) Beta 51%
34 Qwen3.6 35B A3B (Reasoning) Reasoning 50%
35 Gemini 3.1 Pro (Preview) Published configuration 50%
36 Claude Opus 5 Adaptive Reasoning, Max Effort 50%
37 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 50%
38 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 50%
39 Kimi K3 Published configuration 49%
40 Llama 3.1 Instruct 405B Published configuration 49%
41 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 48%
42 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 48%
43 LFM2.5-8B-A1B Published configuration 47%
44 Grok 4.5 high 46%
45 Gemini 3.6 Flash (high) high 46%
46 Gemma 4 E4B (Non-reasoning) Non-reasoning 46%
47 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 45%
48 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 45%
49 KAT Coder Pro V2 Published configuration 44%
50 Nanbeige4.1-3B Published configuration 43%
51 K2 Think V2 Published configuration 41%
52 Magistral Medium 1.2 Published configuration 41%
53 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 40%
54 Gemini 3.5 Flash (medium) medium 40%
55 Gemini 3.5 Flash (high) high 39%
56 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 39%
57 Inkling (xhigh) xhigh 37%
58 JT-35B-Flash Published configuration 37%
59 Nova Micro Published configuration 35%
60 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 34%
61 KAT-Coder-Pro V1 Published configuration 33%
62 Mistral Small 4 (Reasoning) Reasoning 33%
63 Olmo 3.1 32B Think Published configuration 33%
64 ERNIE 4.5 300B A47B Published configuration 33%
65 Nova Premier Published configuration 32%
66 Llama 3.1 Nemotron Instruct 70B Published configuration 31%
67 LFM2 24B A2B Published configuration 30%
68 Olmo 3.1 32B Instruct Published configuration 30%
69 Qwen3.5 4B (Reasoning) Reasoning 28%
70 GPT-5.5 Instant (June 2026) June 2026 27%
71 Hy3 Published configuration 27%
72 Muse Spark Published configuration 27%
73 Gemini 3.5 Flash (minimal) minimal 27%
74 GPT-5.6 Luna (Non-reasoning) Non-reasoning 27%
75 Gemma 4 12B (Non-reasoning) Non-reasoning 27%
76 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 26%
77 Grok 4.3 (Non-reasoning) Non-reasoning 26%
78 Granite 4.0 Micro Published configuration 26%
79 LongCat 2.0 Published configuration 25%
80 MiMo-V2-Flash (Non-reasoning) Non-reasoning 25%
81 Gemma 4 E2B (Non-reasoning) Non-reasoning 25%
82 Phi-4 Mini Instruct Published configuration 24%
83 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 24%
84 Hy3-preview (Non-reasoning) Non-reasoning 24%
85 Command A Published configuration 24%
86 K2-V2 (low) low 24%
87 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 23%
88 Mistral Small 4 (Non-reasoning) Non-reasoning 23%
89 gpt-oss-120b (low) low 22%
90 Llama 4 Scout Published configuration 22%
91 Apertus 70B Instruct Published configuration 21%
92 Cogito v2.1 (Reasoning) Reasoning 21%
93 HyperCLOVA X SEED Think (32B) 32B 21%
94 Doubao Seed Code Published configuration 21%
95 Nova 2.0 Lite (low) low 20%
96 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 20%
97 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 20%
98 Kimi K2.7 Code Published configuration 20%
99 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 20%
100 Phi-4 Published configuration 19%
101 EXAONE 4.5 33B Published configuration 19%
102 Gemma 4 12B (Reasoning) Reasoning 19%
103 Ministral 3 3B Published configuration 19%
104 Gemma 4 26B A4B (Reasoning) Reasoning 19%
105 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 19%
106 Qwen3.5 9B (Reasoning) Reasoning 19%
107 K2-V2 (medium) medium 19%
108 Gemma 4 31B (Reasoning) Reasoning 18%
109 Llama 3.2 Instruct 11B (Vision) Vision 18%
110 Gemini 3.1 Flash-Lite Published configuration 18%
111 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 18%
112 Mistral Medium 3.5 Published configuration 18%
113 Gemma 4 31B (Non-reasoning) Non-reasoning 18%
114 Step3 VL 10B Published configuration 18%
115 Qwen3 Next 80B A3B (Reasoning) Reasoning 18%
116 HyperNova 60B 2605 Published configuration 18%
117 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 17%
118 North Mini Code Published configuration 17%
119 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 17%
120 K-EXAONE (Non-reasoning) Non-reasoning 16%
121 Granite 4.0 350M Published configuration 16%
122 Mistral Large 3 Published configuration 16%
123 Ring-2.6-1T Published configuration 16%
124 Nova 2.0 Lite (Non-reasoning) Non-reasoning 16%
125 Nova 2.0 Omni (low) low 16%
126 Qwen3.6 27B (Non-reasoning) Non-reasoning 16%
127 Step 3.7 Flash Published configuration 16%
128 ERNIE 5.0 Thinking Preview Published configuration 15%
129 Nemotron Cascade 2 30B A3B Published configuration 15%
130 Devstral 2 Published configuration 15%
131 Llama 3.3 Instruct 70B Published configuration 15%
132 Devstral Small 2 Published configuration 15%
133 GPT-5.6 Terra max 15%
134 Qwen3.5 122B A10B (Reasoning) Reasoning 15%
135 Tri-21B-Think Published configuration 15%
136 Jamba Reasoning 3B Published configuration 14%
137 EXAONE 4.0 32B (Reasoning) Reasoning 14%
138 gpt-oss-20b (low) low 13%
139 Trinity Large Thinking Published configuration 13%
140 Agnes 2.5 Pro Alpha Published configuration 13%
141 GPT-5.6 Terra (xhigh) xhigh 13%
142 GPT-5.3 Codex (xhigh) xhigh 13%
143 Hy3-preview (Reasoning) Reasoning 13%
144 INTELLECT-3 Published configuration 13%
145 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 13%
146 GPT-5.6 Sol (medium) medium 13%
147 o3 Published configuration 13%
148 Nova 2.0 Pro Preview (low) low 13%
149 Granite 4.1 8B Published configuration 13%
150 GPT-5.6 Sol (low) low 13%
151 Llama 4 Maverick Published configuration 13%
152 Gemini 2.5 Pro Published configuration 13%
153 GPT-5.6 Terra (high) high 13%
154 LFM2.5-1.2B-Instruct Published configuration 12%
155 Granite 4.0 H 1B Published configuration 12%
156 GPT-5.6 Luna (low) low 12%
157 GPT-5.6 Terra (low) low 12%
158 Solar Pro 3 Published configuration 12%
159 GPT-5.6 Terra (medium) medium 12%
160 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 12%
161 Granite 4.0 H Small Published configuration 12%
162 LFM2 8B A1B Published configuration 12%
163 GPT-5.6 Sol (high) high 12%
164 Solar Open 100B (Reasoning) Reasoning 12%
165 GPT-5.6 Luna (medium) medium 11%
166 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 11%
167 Falcon-H1R-7B Published configuration 11%
168 Qwen3 Omni 30B A3B (Reasoning) Reasoning 11%
169 Nova 2.0 Omni (Non-reasoning) Non-reasoning 11%
170 GPT-5.6 Sol max 11%
171 Mi:dm K 2.5 Pro Published configuration 11%
172 Ministral 3 8B Published configuration 11%
173 GPT-5.6 Sol (xhigh) xhigh 11%
174 Qwen3.5 397B A17B (Reasoning) Reasoning 11%
175 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 11%
176 K-EXAONE (Reasoning) Reasoning 11%
177 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 11%
178 NVIDIA Nemotron 3 Nano 4B Published configuration 11%
179 GPT-5.6 Luna (high) high 10%
180 Ring-flash-2.0 Published configuration 10%
181 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 10%
182 Nova 2.0 Pro Preview (medium) medium 10%
183 Nova 2.0 Lite (high) high 10%
184 Reka Flash 3 Published configuration 10%
185 GPT-5.6 Luna (xhigh) xhigh 10%
186 GPT-5.6 Luna max 10%
187 Ministral 3 14B Published configuration 10%
188 Motif-2-12.7B-Reasoning Published configuration 10%
189 Nova 2.0 Lite (medium) medium 10%
190 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 9%
191 Magistral Small 1.2 Published configuration 9%
192 Molmo2-8B Published configuration 9%
193 Apriel-v1.6-15B-Thinker Published configuration 9%
194 Qwen3 Coder Next Published configuration 9%
195 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 9%
196 DiffusionGemma 26B A4B Published configuration 9%
197 gpt-oss-120b (high) high 9%
198 JT-MINI Published configuration 9%
199 GPT-5.6 Sol (Non-reasoning) Non-reasoning 9%
200 LFM2 2.6B Published configuration 9%
201 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 9%
202 Mercury 2 Published configuration 8%
203 Granite 4.0 H 350M Published configuration 8%
204 K2-V2 (high) high 8%
205 Olmo 3 7B Instruct Published configuration 8%
206 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 8%
207 Nova 2.0 Omni (medium) medium 8%
208 Ling-2.6-1T Published configuration 8%
209 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 8%
210 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 7%
211 Mi:dm K 2.5 Pro Preview Published configuration 7%
212 Qwen3 Next 80B A3B Instruct Published configuration 7%
213 GPT-5.6 Terra (Non-reasoning) Non-reasoning 6%
214 Granite 4.0 1B Published configuration 6%
215 Sarvam 105B (high) high 6%
216 Qwen3.5 Omni Flash Published configuration 6%
217 Jamba 1.7 Large Published configuration 6%
218 LFM2.5-VL-1.6B Published configuration 6%
219 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 6%
220 gpt-oss-20b (high) high 6%
221 Exaone 4.0 1.2B (Reasoning) Reasoning 6%
222 Nex-N2-Pro Published configuration 5%
223 Apertus 8B Instruct Published configuration 5%
224 Olmo 3 7B Think Published configuration 5%
225 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 5%
226 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 5%
227 Granite 4.1 30B Published configuration 5%
228 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 5%
229 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 5%
230 Granite 4.1 3B Published configuration 5%
231 LongCat Flash Lite Published configuration 4%
232 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 4%
233 Tiny Aya Global Published configuration 4%
234 Ling-mini-2.0 Published configuration 4%
235 Ling 2.6 Flash Published configuration 4%
236 Jamba 1.7 Mini Published configuration 3%
237 MiniCPM-V 4.6 1.3B Published configuration 3%
238 LFM2.5-1.2B-Thinking Published configuration 3%
239 Sarvam 30B (high) high 3%
240 Qwen3.5 4B (Non-reasoning) Non-reasoning 3%
241 Qwen3.5 2B (Non-reasoning) Non-reasoning 3%
242 Qwen3 Omni 30B A3B Instruct Published configuration 2%
243 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 1%
244 Qwen3.5 9B (Non-reasoning) Non-reasoning 1%

Showing the top 25 of 244 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

One minus the AA-Omniscience hallucination rate, where incorrect answers are divided by incorrect, partial, and not-attempted outcomes.

How to read the score

The published metric is Non-hallucination rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same omniscienceNonHallucination field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official AA-Omniscience Non-Hallucination Rate resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗