Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

AA-Omniscience Index benchmark: model scores and methodology

A factual-reliability score that rewards correct answers, penalizes hallucinated answers, and leaves abstentions neutral.

Current published leader
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Top score
40
Primary metric
AA-Omniscience Index · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current AA-Omniscience Index model scores

244 configurations · 176 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 244 current configurations with a reported AA-Omniscience Index score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on AA-Omniscience Index

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 40
  2. Gemini 3.1 Pro (Preview) 33
  3. Claude Opus 5 31
  4. Grok 4.5 26
  5. Gemini 3.6 Flash high 24
  6. Gemini 3.5 Flash high 23
  7. GPT-5.6 Sol 22
  8. Kimi K3 18
  9. Muse Spark 1.1 xhigh 18
  10. Grok 4.3 medium 17
  11. Claude Sonnet 5 Adaptive Reasoning, Max Effort 15
  12. Qwen3.7 Max 14
  13. GPT-5.3 Codex xhigh 10
  14. Gemini 3.5 Flash-Lite 7
  15. Muse Spark 4
  16. GLM-5.2 max 4
  17. MiMo-V2.5-Pro 4
  18. GPT-5.5 Instant June 2026 3
  19. Qwen3.6 Plus 3
  20. Qwen3.7 Plus 2
0.0 20.1 40.1

AA-Omniscience Index · higher is better

Current configurations with a reported AA-Omniscience Index value; missing values are omitted.

How to interpret the result

What do AA-Omniscience Index results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher aa-omniscience index is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read AA-Omniscience Index as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationAA-Omniscience Index
1 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Leader Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 40
2 Gemini 3.1 Pro (Preview) Published configuration 33
3 Claude Opus 5 Adaptive Reasoning, Max Effort 31
4 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 30
5 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 28
6 Grok 4.5 high 26
7 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 26
8 Gemini 3.6 Flash (high) high 24
9 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 23
10 Gemini 3.5 Flash (high) high 23
11 GPT-5.6 Sol max 22
12 Gemini 3.5 Flash (medium) medium 22
13 GPT-5.6 Sol (xhigh) xhigh 21
14 GPT-5.6 Sol (high) high 20
15 GPT-5.6 Sol (medium) medium 19
16 Kimi K3 Published configuration 18
17 GPT-5.6 Sol (low) low 18
18 Muse Spark 1.1 (xhigh) xhigh 18
19 Grok 4.3 (medium) medium 17
20 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 15
21 Qwen3.7 Max Published configuration 14
22 Grok 4.3 (low) low 14
23 GPT-5.3 Codex (xhigh) xhigh 10
24 Gemini 3.5 Flash-Lite Published configuration 7
25 Muse Spark Published configuration 4
Show the remaining 219 configurations
#ModelConfigurationAA-Omniscience Index
26 GLM-5.2 (max) max 4
27 MiMo-V2.5-Pro Published configuration 4
28 GPT-5.5 Instant (June 2026) June 2026 3
29 Qwen3.6 Plus Published configuration 3
30 Qwen3.7 Plus Published configuration 2
31 Inkling (xhigh) xhigh 2
32 MiniMax-M3 Published configuration 1
33 Gemini 3.5 Flash (minimal) minimal 1
34 GPT-5.6 Sol (Non-reasoning) Non-reasoning 1
35 GPT-5.6 Terra max 0
36 MiniCPM5-1B (Non-reasoning) Non-reasoning -1
37 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning -1
38 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort -1
39 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort -2
40 GPT-5.6 Terra (xhigh) xhigh -3
41 Command A+ Published configuration -4
42 GPT-5.6 Terra (high) high -4
43 Claude 4.5 Haiku (Reasoning) Reasoning -4
44 G9v3-3B Published configuration -5
45 GPT-5.6 Terra (medium) medium -5
46 GLM-5.2 (Non-reasoning) Non-reasoning -6
47 GPT-5.6 Terra (low) low -7
48 Claude 4.5 Haiku (Non-reasoning) Non-reasoning -8
49 MiMo-V2.5 Published configuration -9
50 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort -10
51 Kimi K2.6 (Non-reasoning) Non-reasoning -10
52 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort -10
53 JT-4.1 Flash 236B A21B Published configuration -10
54 Kimi K2.7 Code Published configuration -11
55 GPT-5.6 Luna max -11
56 GPT-5.6 Luna (xhigh) xhigh -12
57 Qwen3.5 Omni Plus Published configuration -12
58 GPT-5.6 Luna (high) high -12
59 GPT-5.6 Luna (medium) medium -14
60 MiMo-V2-Omni-0327 Published configuration -14
61 Gemini 2.5 Pro Published configuration -14
62 GPT-5.6 Luna (low) low -15
63 o3 Published configuration -15
64 Gemini 3.1 Flash-Lite Published configuration -16
65 MiniCPM5-1B (Reasoning) Reasoning -17
66 Motif 3 (Beta) Beta -17
67 Llama 3.1 Instruct 405B Published configuration -17
68 MiMo-V2-Omni Published configuration -17
69 MiMo-V2-Flash (Feb 2026) Feb 2026 -18
70 Hy3 Published configuration -18
71 Gemma 4 E4B (Reasoning) Reasoning -20
72 Qwen3.6 27B (Reasoning) Reasoning -20
73 Qwen3.6 35B A3B (Reasoning) Reasoning -21
74 KAT Coder Pro V2 Published configuration -22
75 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort -22
76 LongCat 2.0 Published configuration -23
77 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort -23
78 JT-35B-Flash Published configuration -23
79 GPT-5.6 Terra (Non-reasoning) Non-reasoning -23
80 Gemma 4 E2B (Reasoning) Reasoning -25
81 Cogito v2.1 (Reasoning) Reasoning -25
82 GPT-5.6 Luna (Non-reasoning) Non-reasoning -25
83 Magistral Medium 1.2 Published configuration -26
84 Agnes 2.5 Pro Alpha Published configuration -26
85 Nex-N2-Pro Published configuration -28
86 Qwen3.5 397B A17B (Reasoning) Reasoning -30
87 Mistral Small 4 (Reasoning) Reasoning -30
88 DeepSeek V4 Pro (Non-reasoning) Non-reasoning -30
89 Gemma 3 270M Published configuration -31
90 Grok 4.3 (Non-reasoning) Non-reasoning -32
91 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning -33
92 K2 Think V2 Published configuration -34
93 Doubao Seed Code Published configuration -34
94 Hy3-preview (Reasoning) Reasoning -35
95 ERNIE 4.5 300B A47B Published configuration -36
96 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning -36
97 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning -36
98 Nova Premier Published configuration -36
99 Mistral Medium 3.5 Published configuration -36
100 Hy3-preview (Non-reasoning) Non-reasoning -36
101 KAT-Coder-Pro V1 Published configuration -37
102 Step 3.7 Flash Published configuration -38
103 Ring-2.6-1T Published configuration -38
104 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning -38
105 Qwen3.5 0.8B (Reasoning) Reasoning -39
106 LFM2.5-8B-A1B Published configuration -39
107 Mistral Large 3 Published configuration -39
108 Qwen3.5 122B A10B (Reasoning) Reasoning -40
109 Llama 3.1 Nemotron Instruct 70B Published configuration -41
110 Gemma 4 E4B (Non-reasoning) Non-reasoning -42
111 Llama 4 Maverick Published configuration -42
112 Nanbeige4.1-3B Published configuration -42
113 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning -42
114 Qwen3.5 2B (Reasoning) Reasoning -42
115 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning -43
116 Olmo 3.1 32B Think Published configuration -44
117 ERNIE 5.0 Thinking Preview Published configuration -44
118 DeepSeek V4 Flash (Non-reasoning) Non-reasoning -44
119 Trinity Large Thinking Published configuration -44
120 Gemma 4 31B (Reasoning) Reasoning -45
121 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning -46
122 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning -46
123 Nova 2.0 Pro Preview (low) low -46
124 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning -46
125 Devstral 2 Published configuration -46
126 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning -47
127 Nova 2.0 Pro Preview (medium) medium -48
128 Gemma 4 26B A4B (Reasoning) Reasoning -48
129 K2-V2 (low) low -48
130 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning -48
131 Mistral Small 4 (Non-reasoning) Non-reasoning -48
132 Command A Published configuration -48
133 MiMo-V2-Flash (Non-reasoning) Non-reasoning -48
134 Nova Micro Published configuration -49
135 North Mini Code Published configuration -49
136 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning -50
137 K2-V2 (medium) medium -50
138 Nova 2.0 Omni (low) low -50
139 gpt-oss-120b (high) high -50
140 gpt-oss-120b (low) low -50
141 Qwen3 Next 80B A3B (Reasoning) Reasoning -51
142 Ling-2.6-1T Published configuration -51
143 INTELLECT-3 Published configuration -51
144 Nova 2.0 Lite (low) low -51
145 Olmo 3.1 32B Instruct Published configuration -51
146 Gemma 4 31B (Non-reasoning) Non-reasoning -51
147 EXAONE 4.5 33B Published configuration -51
148 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning -52
149 Qwen3.5 4B (Reasoning) Reasoning -52
150 Gemma 4 12B (Reasoning) Reasoning -52
151 Llama 3.3 Instruct 70B Published configuration -52
152 Mercury 2 Published configuration -52
153 Llama 4 Scout Published configuration -52
154 Nemotron Cascade 2 30B A3B Published configuration -52
155 Qwen3.5 9B (Reasoning) Reasoning -53
156 Gemma 4 12B (Non-reasoning) Non-reasoning -53
157 Qwen3.6 27B (Non-reasoning) Non-reasoning -53
158 HyperCLOVA X SEED Think (32B) 32B -53
159 Solar Pro 3 Published configuration -54
160 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning -54
161 Solar Open 100B (Reasoning) Reasoning -54
162 Nova 2.0 Lite (high) high -54
163 Apertus 70B Instruct Published configuration -55
164 Jamba 1.7 Large Published configuration -56
165 Nova 2.0 Lite (medium) medium -56
166 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration -56
167 HyperNova 60B 2605 Published configuration -56
168 K2-V2 (high) high -57
169 Phi-4 Published configuration -57
170 Devstral Small 2 Published configuration -57
171 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning -57
172 Mi:dm K 2.5 Pro Published configuration -57
173 Nova 2.0 Omni (medium) medium -58
174 K-EXAONE (Reasoning) Reasoning -58
175 Apriel-v1.6-15B-Thinker Published configuration -59
176 Ring-flash-2.0 Published configuration -59
177 Step3 VL 10B Published configuration -59
178 DiffusionGemma 26B A4B Published configuration -59
179 LFM2 24B A2B Published configuration -59
180 Nova 2.0 Lite (Non-reasoning) Non-reasoning -59
181 Qwen3 Next 80B A3B Instruct Published configuration -59
182 Sarvam 105B (high) high -60
183 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning -60
184 Granite 4.0 Micro Published configuration -60
185 gpt-oss-20b (low) low -60
186 Qwen3 Coder Next Published configuration -61
187 Qwen3 Omni 30B A3B (Reasoning) Reasoning -61
188 EXAONE 4.0 32B (Reasoning) Reasoning -61
189 Granite 4.0 H Small Published configuration -61
190 Phi-4 Mini Instruct Published configuration -61
191 K-EXAONE (Non-reasoning) Non-reasoning -61
192 Motif-2-12.7B-Reasoning Published configuration -61
193 Gemma 4 E2B (Non-reasoning) Non-reasoning -61
194 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning -62
195 Mi:dm K 2.5 Pro Preview Published configuration -62
196 Gemma 4 26B A4B (Non-reasoning) Non-reasoning -62
197 Falcon-H1R-7B Published configuration -62
198 EXAONE 4.0 32B (Non-reasoning) Non-reasoning -62
199 Llama 3.2 Instruct 11B (Vision) Vision -63
200 Tri-21B-Think Published configuration -63
201 gpt-oss-20b (high) high -64
202 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning -64
203 Nova 2.0 Omni (Non-reasoning) Non-reasoning -64
204 JT-MINI Published configuration -65
205 Granite 4.1 8B Published configuration -65
206 Reka Flash 3 Published configuration -65
207 Magistral Small 1.2 Published configuration -65
208 Qwen3.5 Omni Flash Published configuration -66
209 Ling 2.6 Flash Published configuration -66
210 Ministral 3 14B Published configuration -67
211 Ministral 3 3B Published configuration -68
212 Granite 4.1 30B Published configuration -68
213 Ministral 3 8B Published configuration -68
214 Molmo2-8B Published configuration -69
215 Qwen3 Omni 30B A3B Instruct Published configuration -69
216 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning -69
217 LongCat Flash Lite Published configuration -70
218 Qwen3.5 9B (Non-reasoning) Non-reasoning -71
219 NVIDIA Nemotron 3 Nano 4B Published configuration -72
220 Sarvam 30B (high) high -72
221 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning -73
222 Jamba Reasoning 3B Published configuration -74
223 Olmo 3 7B Think Published configuration -74
224 Jamba 1.7 Mini Published configuration -74
225 Apertus 8B Instruct Published configuration -75
226 Qwen3.5 4B (Non-reasoning) Non-reasoning -75
227 LFM2 8B A1B Published configuration -75
228 LFM2.5-1.2B-Instruct Published configuration -76
229 LFM2 2.6B Published configuration -77
230 Granite 4.1 3B Published configuration -77
231 Olmo 3 7B Instruct Published configuration -78
232 Granite 4.0 H 1B Published configuration -78
233 Ling-mini-2.0 Published configuration -79
234 Granite 4.0 350M Published configuration -79
235 Granite 4.0 1B Published configuration -82
236 Exaone 4.0 1.2B (Reasoning) Reasoning -82
237 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning -83
238 Qwen3.5 2B (Non-reasoning) Non-reasoning -83
239 LFM2.5-1.2B-Thinking Published configuration -84
240 LFM2.5-VL-1.6B Published configuration -84
241 MiniCPM-V 4.6 1.3B Published configuration -85
242 Tiny Aya Global Published configuration -85
243 Granite 4.0 H 350M Published configuration -85
244 Qwen3.5 0.8B (Non-reasoning) Non-reasoning -89

Showing the top 25 of 244 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A factual-reliability score that rewards correct answers, penalizes hallucinated answers, and leaves abstentions neutral.

How to read the score

The published metric is AA-Omniscience Index. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same omniscience field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official AA-Omniscience Index resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗