Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

AA-LCR benchmark: model scores and methodology

Artificial Analysis's long-context benchmark for answering questions that require reasoning across multiple documents.

Current published leader
Kimi K3 · Published configuration
Top score
75%
Primary metric
Accuracy · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current AA-LCR model scores

246 configurations · 178 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 246 current configurations with a reported AA-LCR score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on AA-LCR

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. These results sit in a narrow band, so the scale below is zoomed — read each mark against the labelled axis, not the left edge. Marks are placed on the full source precision, so two models sharing a rounded label can still sit at different points. Score labels reproduce Artificial Analysis's rounded display value.

  1. Kimi K3 75%
  2. GPT-5.3 Codex xhigh 74%
  3. GPT-5.6 Luna 74%
  4. GPT-5.6 Terra 74%
  5. KAT-Coder-Pro V1 74%
  6. MiniMax-M3 74%
  7. GPT-5.6 Sol 74%
  8. MiMo-V2.5-Pro 73%
  9. Gemini 3.1 Pro (Preview) 73%
  10. GLM-5.2 max 71%
  11. Gemini 3.5 Flash medium 71%
  12. Claude Sonnet 5 Adaptive Reasoning, Max Effort 71%
  13. Claude 4.5 Haiku Reasoning 70%
  14. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 70%
  15. Claude Opus 5 70%
  16. Gemini 3.6 Flash high 70%
  17. Muse Spark 70%
  18. Qwen3.6 Plus 70%
  19. o3 69%
  20. Qwen3.7 Max 69%
67.0 71.2 75.3

Accuracy · higher is better · zoomed scale

Current configurations with a reported AA-LCR value; missing values are omitted.

How to interpret the result

What do AA-LCR results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher accuracy is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read AA-LCR as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationAccuracy
1 Kimi K3Leader Published configuration 75%
2 GPT-5.3 Codex (xhigh) xhigh 74%
3 GPT-5.6 Luna max 74%
4 GPT-5.6 Terra max 74%
5 KAT-Coder-Pro V1 Published configuration 74%
6 MiniMax-M3 Published configuration 74%
7 GPT-5.6 Sol max 74%
8 MiMo-V2.5-Pro Published configuration 73%
9 Gemini 3.1 Pro (Preview) Published configuration 73%
10 GPT-5.6 Terra (high) high 72%
11 GLM-5.2 (max) max 71%
12 GPT-5.6 Terra (xhigh) xhigh 71%
13 GPT-5.6 Sol (xhigh) xhigh 71%
14 Gemini 3.5 Flash (medium) medium 71%
15 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 71%
16 Claude 4.5 Haiku (Reasoning) Reasoning 70%
17 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 70%
18 Claude Opus 5 Adaptive Reasoning, Max Effort 70%
19 GPT-5.6 Luna (xhigh) xhigh 70%
20 Gemini 3.6 Flash (high) high 70%
21 Muse Spark Published configuration 70%
22 Qwen3.6 Plus Published configuration 70%
23 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 69%
24 Gemini 3.5 Flash (high) high 69%
25 o3 Published configuration 69%
Show the remaining 221 configurations
#ModelConfigurationAccuracy
26 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 69%
27 GPT-5.6 Luna (high) high 69%
28 Qwen3.7 Max Published configuration 69%
29 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 69%
30 GPT-5.6 Sol (medium) medium 69%
31 Qwen3.6 27B (Reasoning) Reasoning 69%
32 GPT-5.6 Sol (high) high 68%
33 GPT-5.6 Terra (medium) medium 68%
34 GPT-5.6 Sol (low) low 68%
35 Grok 4.5 high 68%
36 Nex-N2-Pro Published configuration 68%
37 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 67%
38 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 67%
39 Hy3 Published configuration 67%
40 MiMo-V2-Omni Published configuration 67%
41 Qwen3.5 122B A10B (Reasoning) Reasoning 67%
42 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 66%
43 Kimi K2.7 Code Published configuration 66%
44 GPT-5.6 Luna (medium) medium 66%
45 Gemini 2.5 Pro Published configuration 66%
46 KAT Coder Pro V2 Published configuration 66%
47 Qwen3.5 397B A17B (Reasoning) Reasoning 66%
48 Doubao Seed Code Published configuration 65%
49 Gemini 3.1 Flash-Lite Published configuration 65%
50 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 65%
51 Grok 4.3 (medium) medium 65%
52 Qwen3.7 Plus Published configuration 65%
53 GPT-5.6 Terra (low) low 64%
54 MiMo-V2-Flash (Feb 2026) Feb 2026 64%
55 Ring-2.6-1T Published configuration 64%
56 GPT-5.5 Instant (June 2026) June 2026 64%
57 Grok 4.3 (low) low 64%
58 Agnes 2.5 Pro Alpha Published configuration 64%
59 MiMo-V2-Omni-0327 Published configuration 64%
60 Qwen3.6 35B A3B (Reasoning) Reasoning 64%
61 Step 3.7 Flash Published configuration 64%
62 Inkling (xhigh) xhigh 63%
63 Muse Spark 1.1 (xhigh) xhigh 63%
64 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 63%
65 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 63%
66 MiMo-V2.5 Published configuration 63%
67 Gemini 3.5 Flash-Lite Published configuration 62%
68 Gemma 4 31B (Reasoning) Reasoning 62%
69 Nova 2.0 Pro Preview (low) low 62%
70 Motif 3 (Beta) Beta 61%
71 Mistral Medium 3.5 Published configuration 61%
72 Qwen3 Next 80B A3B (Reasoning) Reasoning 60%
73 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 60%
74 GPT-5.6 Luna (low) low 59%
75 Qwen3.5 9B (Reasoning) Reasoning 59%
76 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 59%
77 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 59%
78 JT-4.1 Flash 236B A21B Published configuration 59%
79 Nova 2.0 Lite (medium) medium 58%
80 LongCat 2.0 Published configuration 58%
81 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 58%
82 Kimi K2.6 (Non-reasoning) Non-reasoning 58%
83 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 57%
84 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 56%
85 Gemma 4 26B A4B (Reasoning) Reasoning 56%
86 K-EXAONE (Reasoning) Reasoning 56%
87 Qwen3.5 4B (Reasoning) Reasoning 56%
88 GPT-5.6 Sol (Non-reasoning) Non-reasoning 55%
89 Gemma 4 12B (Reasoning) Reasoning 55%
90 JT-35B-Flash Published configuration 55%
91 Nova 2.0 Lite (high) high 55%
92 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 55%
93 Qwen3.6 27B (Non-reasoning) Non-reasoning 55%
94 Hy3-preview (Reasoning) Reasoning 55%
95 Nova 2.0 Pro Preview (medium) medium 54%
96 Nova 2.0 Omni (medium) medium 54%
97 Gemini 3.5 Flash (minimal) minimal 53%
98 K2 Think V2 Published configuration 53%
99 Qwen3.5 Omni Plus Published configuration 53%
100 Nova 2.0 Lite (low) low 52%
101 Magistral Medium 1.2 Published configuration 51%
102 Qwen3 Next 80B A3B Instruct Published configuration 51%
103 Nova 2.0 Omni (low) low 51%
104 gpt-oss-120b (high) high 51%
105 Apriel-v1.6-15B-Thinker Published configuration 50%
106 GPT-5.6 Terra (Non-reasoning) Non-reasoning 50%
107 EXAONE 4.5 33B Published configuration 49%
108 K-EXAONE (Non-reasoning) Non-reasoning 47%
109 Command A+ Published configuration 46%
110 Llama 4 Maverick Published configuration 46%
111 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 45%
112 Mistral Small 4 (Reasoning) Reasoning 45%
113 Qwen3.5 Omni Flash Published configuration 44%
114 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 44%
115 gpt-oss-120b (low) low 44%
116 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 40%
117 Qwen3 Coder Next Published configuration 40%
118 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 40%
119 Qwen3.5 9B (Non-reasoning) Non-reasoning 38%
120 GLM-5.2 (Non-reasoning) Non-reasoning 37%
121 GPT-5.6 Luna (Non-reasoning) Non-reasoning 36%
122 Mercury 2 Published configuration 36%
123 Gemma 4 31B (Non-reasoning) Non-reasoning 36%
124 Solar Open 100B (Reasoning) Reasoning 36%
125 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 36%
126 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 35%
127 G9v3-3B Published configuration 35%
128 Ling-2.6-1T Published configuration 35%
129 Mistral Large 3 Published configuration 35%
130 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 34%
131 Nemotron Cascade 2 30B A3B Published configuration 34%
132 Hy3-preview (Non-reasoning) Non-reasoning 34%
133 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 34%
134 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 33%
135 K2-V2 (high) high 33%
136 Trinity Large Thinking Published configuration 33%
137 INTELLECT-3 Published configuration 32%
138 North Mini Code Published configuration 32%
139 HyperNova 60B 2605 Published configuration 32%
140 MiMo-V2-Flash (Non-reasoning) Non-reasoning 31%
141 gpt-oss-20b (low) low 31%
142 Gemma 4 12B (Non-reasoning) Non-reasoning 31%
143 Gemma 4 E4B (Reasoning) Reasoning 31%
144 gpt-oss-20b (high) high 31%
145 Devstral 2 Published configuration 30%
146 Nova Premier Published configuration 30%
147 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 28%
148 Qwen3.5 4B (Non-reasoning) Non-reasoning 28%
149 K2-V2 (medium) medium 28%
150 Solar Pro 3 Published configuration 27%
151 Llama 4 Scout Published configuration 26%
152 Kimi Linear 48B A3B Instruct Published configuration 26%
153 LongCat Flash Lite Published configuration 26%
154 Ling 2.6 Flash Published configuration 25%
155 Grok 4.3 (Non-reasoning) Non-reasoning 25%
156 Llama 3.1 Instruct 405B Published configuration 24%
157 Devstral Small 2 Published configuration 24%
158 Ministral 3 8B Published configuration 24%
159 Qwen3.5 2B (Reasoning) Reasoning 24%
160 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 23%
161 Nova 2.0 Omni (Non-reasoning) Non-reasoning 22%
162 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 22%
163 Ministral 3 14B Published configuration 22%
164 Cogito v2.1 (Reasoning) Reasoning 22%
165 Mistral Small 4 (Non-reasoning) Non-reasoning 21%
166 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 21%
167 Ring-flash-2.0 Published configuration 21%
168 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 21%
169 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 20%
170 K2-V2 (low) low 19%
171 Granite 4.1 30B Published configuration 19%
172 Command A Published configuration 18%
173 Gemma 4 E4B (Non-reasoning) Non-reasoning 18%
174 Nova 2.0 Lite (Non-reasoning) Non-reasoning 18%
175 Jamba 1.7 Large Published configuration 17%
176 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 17%
177 NVIDIA Nemotron 3 Nano 4B Published configuration 17%
178 Magistral Small 1.2 Published configuration 16%
179 Gemma 4 E2B (Non-reasoning) Non-reasoning 15%
180 Gemma 4 E2B (Reasoning) Reasoning 15%
181 Llama 3.3 Instruct 70B Published configuration 15%
182 DiffusionGemma 26B A4B Published configuration 14%
183 EXAONE 4.0 32B (Reasoning) Reasoning 14%
184 Phi-4 Mini Instruct Published configuration 14%
185 Qwen3.5 2B (Non-reasoning) Non-reasoning 14%
186 Motif-2-12.7B-Reasoning Published configuration 13%
187 Jamba 1.7 Mini Published configuration 13%
188 Granite 4.1 8B Published configuration 12%
189 HyperCLOVA X SEED Think (32B) 32B 12%
190 JT-MINI Published configuration 12%
191 Llama 3.2 Instruct 11B (Vision) Vision 12%
192 Ministral 3 3B Published configuration 12%
193 Mi:dm K 2.5 Pro Preview Published configuration 11%
194 Tri-21B-Think Published configuration 11%
195 Nova Micro Published configuration 10%
196 Granite 4.0 H Small Published configuration 9%
197 Mi:dm K 2.5 Pro Published configuration 9%
198 Falcon-H1R-7B Published configuration 9%
199 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 8%
200 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 7%
201 Jamba Reasoning 3B Published configuration 7%
202 Llama 3.1 Nemotron Instruct 70B Published configuration 7%
203 ERNIE 5.0 Thinking Preview Published configuration 7%
204 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 7%
205 Ling-mini-2.0 Published configuration 7%
206 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 7%
207 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 7%
208 Granite 4.0 H 1B Published configuration 6%
209 MiniCPM-V 4.6 1.3B Published configuration 6%
210 Qwen3.5 0.8B (Reasoning) Reasoning 5%
211 MiniCPM5-1B (Non-reasoning) Non-reasoning 5%
212 Granite 4.0 1B Published configuration 4%
213 Granite 4.0 Micro Published configuration 4%
214 MiniCPM5-1B (Reasoning) Reasoning 4%
215 Granite 4.1 3B Published configuration 3%
216 ERNIE 4.5 300B A47B Published configuration 2%
217 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 2%
218 Apertus 70B Instruct Published configuration 0%
219 Apertus 8B Instruct Published configuration 0%
220 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 0%
221 Exaone 4.0 1.2B (Reasoning) Reasoning 0%
222 Gemma 3 270M Published configuration 0%
223 Granite 4.0 350M Published configuration 0%
224 Granite 4.0 H 350M Published configuration 0%
225 LFM2 2.6B Published configuration 0%
226 LFM2 24B A2B Published configuration 0%
227 LFM2 8B A1B Published configuration 0%
228 LFM2.5-1.2B-Instruct Published configuration 0%
229 LFM2.5-1.2B-Thinking Published configuration 0%
230 LFM2.5-8B-A1B Published configuration 0%
231 LFM2.5-VL-1.6B Published configuration 0%
232 Molmo 7B-D Published configuration 0%
233 Molmo2-8B Published configuration 0%
234 Nanbeige4.1-3B Published configuration 0%
235 Olmo 3 7B Instruct Published configuration 0%
236 Olmo 3 7B Think Published configuration 0%
237 Olmo 3.1 32B Instruct Published configuration 0%
238 Olmo 3.1 32B Think Published configuration 0%
239 Phi-4 Published configuration 0%
240 Qwen3 Omni 30B A3B (Reasoning) Reasoning 0%
241 Qwen3 Omni 30B A3B Instruct Published configuration 0%
242 Reka Flash 3 Published configuration 0%
243 Sarvam 105B (high) high 0%
244 Sarvam 30B (high) high 0%
245 Step3 VL 10B Published configuration 0%
246 Tiny Aya Global Published configuration 0%

Showing the top 25 of 246 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

Artificial Analysis's long-context benchmark for answering questions that require reasoning across multiple documents.

How to read the score

The published metric is Accuracy. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same lcr field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official AA-LCR resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗