Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of questions answered correctly over long documents, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
393 configurations of 233 models
Unit
% of questions
Licence
Artificial Analysis commercial data licence

Long-context reasoning (AA-LCR): Gemini 4 Argon

Top 15 of 233 results · % of questions, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Kimi K3 (max reasoning)Moonshot AI 88.7%
  2. 2 MiMo-V2.6-ProXiaomi 86.3%
  3. 3 Claude Fable 5.1 (max reasoning)Anthropic 85.3%
  4. 4 Claude Opus 5.5 (max reasoning)Anthropic 84.7%
  5. 5 GPT-5.5 (high reasoning)OpenAI 84.3%
  6. 6 DeepSeek-V4.1-Flash (max reasoning)DeepSeek 84.0%
  7. 6 Gemini 3.8 Flash (medium reasoning)Google 84.0%
  8. 6 GPT-5.6 Sol (max reasoning)OpenAI 84.0%
  9. 6 GPT-6.1 Sol (low reasoning)OpenAI 84.0%
  10. 10 GPT-5.6 Luna (max reasoning)OpenAI 83.7%
  11. 10 GPT-6 Sol (high reasoning)OpenAI 83.7%
  12. 12 Muse Glimmer (high reasoning)Meta 83.3%
  13. 12 GPT-5.3-Codex (extra-high reasoning)OpenAI 83.3%
  14. 12 GPT-6 Luna (max reasoning)OpenAI 83.3%
  15. 15 Gemini 3.7 Flash (medium reasoning)Google 83.0%
  16. 43 Gemini 4 Argon (high reasoning)Google 79.7%

Full results

Artificial Analysis: Long-context reasoning (AA-LCR), % of questions, higher is better
#ModelLong-context reasoning (AA-LCR)
% of questions, higher is better
Price
$ per million tokens, in / out
1 Kimi K3 (max reasoning)Moonshot AI · best of 2 settings
88.7%
$3 / $15
2 MiMo-V2.6-ProXiaomi
86.3%
$0.43 / $0.87
3 Claude Fable 5.1 (max reasoning)Anthropic · best of 5 settings
85.3%
$10 / $50
4 Claude Opus 5.5 (max reasoning)Anthropic · best of 5 settings
84.7%
$4 / $20
5 GPT-5.5 (high reasoning)OpenAI · best of 5 settings
84.3%
$5 / $30
6 DeepSeek-V4.1-Flash (max reasoning)DeepSeek · best of 2 settings
84.0%
$0.30 / $1.20
6 Gemini 3.8 Flash (medium reasoning)Google · best of 3 settings
84.0%
$1.50 / $7.50
6 GPT-5.6 Sol (max reasoning)OpenAI · best of 6 settings
84.0%
$4 / $20
6 GPT-6.1 Sol (low reasoning)OpenAI · best of 5 settings
84.0%
$2 / $10
10 GPT-5.6 Luna (max reasoning)OpenAI · best of 6 settings
83.7%
$0.20 / $1.20
10 GPT-6 Sol (high reasoning)OpenAI · best of 6 settings
83.7%
$2 / $10
12 Muse Glimmer (high reasoning)Meta
83.3%
–
12 GPT-5.3-Codex (extra-high reasoning)OpenAI
83.3%
$1.75 / $14
12 GPT-6 Luna (max reasoning)OpenAI · best of 6 settings
83.3%
$0.10 / $0.50
15 Gemini 3.7 Flash (medium reasoning)Google · best of 3 settings
83.0%
$1.50 / $7.50
15 Muse Spark 1.3 (max reasoning)Meta · best of 2 settings
83.0%
$1.25 / $4.25
15 MiniMax-M3MiniMax
83.0%
$0.30 / $1.20
15 GPT-5.6 Terra (max reasoning)OpenAI · best of 6 settings
83.0%
$2 / $12
19 Claude Haiku 5.5 (max reasoning)Anthropic · best of 5 settings
82.7%
$0.10 / $0.50
19 Claude Sonnet 5.5 (max reasoning)Anthropic · best of 5 settings
82.7%
$2 / $10
19 GPT-5.2 (extra-high reasoning)OpenAI · best of 3 settings
82.7%
$1.75 / $14
22 Claude Fable 5 (max reasoning)Anthropic
82.3%
$10 / $50
22 GPT-5.2-Codex (extra-high reasoning)OpenAI
82.3%
$1.75 / $14
24 Qwen3.8-27B (extra-high reasoning)Alibaba · best of 4 settings
82.0%
$0.50 / $3
24 Claude Opus 5 (medium reasoning)Anthropic · best of 5 settings
82.0%
$5 / $25
24 Claude Sonnet 5 (max reasoning)Anthropic · best of 6 settings
82.0%
$2 / $10
24 Gemini 3.1 Pro PreviewGoogle
82.0%
$2 / $12
24 GPT-5.4 (extra-high reasoning)OpenAI · best of 3 settings
82.0%
$2.50 / $15
29 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
81.3%
$0.44 / $1.32
29 Mistral Large 4Mistral AI
81.3%
$1.36 / $4.18
31 Kimi K2.6Moonshot AI · best of 2 settings
81.0%
$0.95 / $4
31 Grok 4.6 (medium reasoning)xAI · best of 4 settings
81.0%
$2 / $6
33 Qwen3.6-Max-PreviewAlibaba
80.7%
$1.03 / $6.16
33 GPT-6 Astra (max reasoning)OpenAI · best of 5 settings
80.7%
$10 / $50
35 Qwen3.8-2.4T-A95BAlibaba
80.3%
$2 / $6
35 Qwen3.8-Max (0902)Alibaba
80.3%
$2 / $6
35 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek · best of 2 settings
80.3%
$1.32 / $3.96
38 Claude Sonnet 4.6 (max reasoning)Anthropic · best of 2 settings
80.0%
$3 / $15
38 Gemini 3.6 Flash (high reasoning)Google
80.0%
$1.50 / $7.50
38 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
80.0%
$1.25 / $10
38 Grok 4.7 (low reasoning)xAI · best of 3 settings
80.0%
$2 / $6
38 GLM-5.3-FlashZ.ai
80.0%
$0.15 / $0.50
43 Qwen3.8-Flash-NextAlibaba
79.7%
–
43 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
79.7%
$0.14 / $0.28
43 Gemini 4 Argon (high reasoning)Google
79.7%
–
43 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
79.7%
$0.43 / $0.87
43 GLM-5.3 (max reasoning)Z.ai · best of 2 settings
79.7%
$1.40 / $4.40
48 Kimi K2.7 CodeMoonshot AI
79.3%
$0.95 / $4
48 Grok 4.5 (high reasoning)xAI
79.3%
$2 / $6
50 Qwen3.7-MaxAlibaba
79.0%
$1.48 / $4.42
50 Muse Spark 1.2 (extra-high reasoning)Meta
79.0%
$1.25 / $4.25
50 Hy3Tencent
79.0%
$0.14 / $0.58
53 Claude Opus 4.7 (max reasoning)Anthropic · best of 2 settings
78.7%
$5 / $25
54 Qwen3.6-PlusAlibaba
78.3%
$0.33 / $1.95
54 Qwen3.8-Max (0803)Alibaba
78.3%
–
54 MiniMax-M2.7MiniMax
78.3%
$0.30 / $1.20
54 GLM-5.2 (max reasoning)Z.ai · best of 2 settings
78.3%
$1.40 / $4.40
58 GPT-5 (high reasoning)OpenAI · best of 2 settings
78.2%
$1.25 / $10
59 Claude Opus 4.6 (max reasoning)Anthropic · best of 2 settings
78.0%
$5 / $25
59 Gemini 3 Flash Preview (reasoning on)Google · best of 2 settings
78.0%
$0.50 / $3
59 Muse SparkMeta
78.0%
–
59 Kimi K2.5Moonshot AI · best of 2 settings
78.0%
$0.57 / $2.85
63 Qwen3.5-27BAlibaba · best of 2 settings
77.7%
$0.27 / $2.16
63 Claude Opus 4.8 (max reasoning)Anthropic
77.7%
$5 / $25
63 Muse Spark 1.1 (extra-high reasoning)Meta
77.7%
$1.25 / $4.25
66 Qwen3.5-397B-A17BAlibaba · best of 2 settings
77.3%
$0.55 / $3.50
66 Qwen3.6-27BAlibaba · best of 2 settings
77.3%
$0.30 / $3.20
66 Claude Opus 4.5 (reasoning on)Anthropic · best of 2 settings
77.3%
$5 / $25
66 Inkling (extra-high reasoning)Thinking Machines
77.3%
$0.95 / $4.05
70 GPT-5.4 mini (extra-high reasoning)OpenAI · best of 3 settings
77.0%
$0.75 / $4.50
71 GPT-5.4 nano (extra-high reasoning)OpenAI · best of 3 settings
76.7%
$0.20 / $1.25
72 Qwen3.5-122B-A10BAlibaba · best of 2 settings
76.3%
$0.26 / $2.08
73 Claude Opus 4.1 (reasoning on)Anthropic
76.0%
$15 / $75
73 Gemini 3.5 Flash-LiteGoogle
76.0%
$0.30 / $2.50
73 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
76.0%
–
76 Inkling SmallThinking Machines
75.7%
$0.45 / $1.20
76 GLM-5Z.ai · best of 2 settings
75.7%
$0.95 / $2.55
78 Grok 4.3 (medium reasoning)xAI · best of 4 settings
75.0%
$1.25 / $2.50
79 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek · best of 3 settings
74.7%
$1.42 / $2.83
79 o3OpenAI
74.7%
$2 / $8
79 Grok Build 0.1xAI
74.7%
$1 / $2
82 Qwen3-Max-ThinkingAlibaba
74.3%
$0.78 / $3.90
82 Claude Haiku 4.5 (reasoning on)Anthropic · best of 2 settings
74.3%
$1 / $5
82 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek · best of 3 settings
74.3%
$0.14 / $0.28
82 Gemini 3.1 Flash-Lite PreviewGoogle
74.3%
$0.25 / $1.50
82 Gemini 3.5 Flash (medium reasoning)Google · best of 3 settings
74.3%
$1.50 / $9
82 MiMo-V2.6-FlashXiaomi
74.3%
$0.14 / $0.28
88 Solar Pro 4Upstage
74.0%
$0.09 / $0.36
88 Grok 4.1 Fast (reasoning)xAI
74.0%
–
90 Grok 4 Fast (reasoning)xAI
73.7%
–
90 GLM-5.1Z.ai · best of 2 settings
73.7%
$1.38 / $4.40
92 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
73.3%
$0.30 / $0.96
92 MiniMax-M2.5MiniMax
73.3%
$0.30 / $1.20
94 Qwen3.7-PlusAlibaba
73.0%
$0.32 / $1.28
94 MiMo-V2.5Xiaomi
73.0%
$0.17 / $0.34
96 Claude Sonnet 4.5 (reasoning on)Anthropic · best of 2 settings
72.3%
$3 / $15
96 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
72.3%
$0.27 / $0.41
96 GPT-5 mini (high reasoning)OpenAI · best of 3 settings
72.3%
$0.25 / $2
99 Qwen3-235B-A22B-Thinking-2507Alibaba
72.0%
$0.30 / $3
99 Qwen3.5-35B-A3BAlibaba · best of 2 settings
72.0%
$0.16 / $1.30
99 Kimi K2 ThinkingMoonshot AI
72.0%
$0.60 / $2.50
102 Qwen3.6-35B-A3BAlibaba · best of 2 settings
71.7%
$0.10 / $1
102 GPT-5-Codex (high reasoning)OpenAI
71.7%
–
102 GLM-5-TurboZ.ai
71.7%
$1.20 / $4
105 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google · best of 2 settings
71.0%
–
105 GLM-4.7Z.ai · best of 2 settings
71.0%
$0.54 / $1.98
107 Claude Sonnet 4 (reasoning on)Anthropic · best of 2 settings
70.3%
$3 / $15
107 GLM-5V-TurboZ.ai
70.3%
$1.20 / $4
109 Qwen3.5-9BAlibaba · best of 2 settings
70.0%
$0.10 / $0.15
109 DeepSeek-V3.2-SpecialeDeepSeek
70.0%
–
109 GPT-5.5 Instant (2026-06-26)OpenAI · best of 2 settings
70.0%
–
112 Gemma 4 31BGoogle · best of 2 settings
69.7%
$0.14 / $0.40
113 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek · best of 2 settings
69.3%
$0.27 / $1
113 Mistral Medium 3.5Mistral AI
69.3%
$1.50 / $7.50
113 GPT-5.1-Codex (high reasoning)OpenAI
69.3%
$1.25 / $10
116 Gemini 2.5 ProGoogle
69.0%
$1.25 / $10
116 Grok 4.20xAI · best of 2 settings
69.0%
$1.25 / $2.50
118 GPT-4.1OpenAI
68.3%
$2 / $8
119 Grok 4xAI
68.0%
–
120 MiniMax-M2.1MiniMax
67.7%
$0.30 / $1.20
120 Grok 4.20 (0309, reasoning)xAI
67.7%
–
122 GPT-5.1-Codex-Mini (high reasoning)OpenAI
66.7%
$0.25 / $2
123 Gemma 4 26B A4BGoogle · best of 2 settings
65.7%
$0.10 / $0.30
123 Nemotron 3 Super 120B A12BNVIDIA
65.7%
$0.085 / $0.40
125 Gemini 2.5 Flash (reasoning on)Google · best of 2 settings
65.3%
$0.30 / $2.50
126 GPT-5 ChatOpenAI
65.0%
–
126 o1OpenAI
65.0%
$15 / $60
128 Gemini 2.5 Flash-Lite Preview (09-2025, reasoning on)Google · best of 2 settings
64.7%
–
128 Hy3 Preview (reasoning on)Tencent · best of 2 settings
64.7%
$0.18 / $0.60
130 MiniMax-M2MiniMax
64.3%
$0.30 / $1.20
131 Qwen3.5-Omni-PlusAlibaba
63.7%
–
131 Qwen3-Next-80B-A3B-ThinkingAlibaba
63.7%
$0.15 / $1.20
131 Qwen3-VL-235B-A22B-ThinkingAlibaba
63.7%
$0.40 / $4
131 Gemma 4 12BGoogle · best of 2 settings
63.7%
–
135 Qwen3.5-4B (reasoning on)Alibaba · best of 2 settings
63.0%
–
136 Qwen3-Max-Thinking-PreviewAlibaba
62.0%
–
137 Qwen3-30B-A3B-Thinking-2507Alibaba
61.3%
$0.20 / $2.40
138 o4-mini (high reasoning)OpenAI
61.0%
$1.10 / $4.40
139 Nova 2 Lite (high reasoning)Amazon · best of 4 settings
60.3%
$0.30 / $2.50
140 DeepSeek-R1DeepSeek
57.7%
$0.70 / $2.50
141 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
56.7%
$0.55 / $1.65
142 DeepSeek-R1-0528DeepSeek
55.7%
$0.50 / $2.18
142 Gemini 2.5 Flash-Lite (reasoning on)Google · best of 2 settings
55.7%
$0.10 / $0.40
144 Qwen3-VL-32B-ThinkingAlibaba
55.3%
–
145 GLM-4.6 (reasoning on)Z.ai · best of 2 settings
54.0%
$0.50 / $2
146 Magistral Medium 1.2Mistral AI
53.0%
–
146 Kimi K2 (0905)Moonshot AI
53.0%
$0.60 / $2.50
146 Kimi K2 (0711)Moonshot AI
53.0%
$0.57 / $2.30
146 Grok Code Fast 1xAI
53.0%
–
150 Qwen3-Next-80B-A3B-InstructAlibaba
52.7%
$0.10 / $1.10
150 GLM-4.5Z.ai
52.7%
$0.60 / $2.20
152 Qwen3.5-Omni-FlashAlibaba
52.0%
–
152 gpt-oss-120b (high reasoning)OpenAI · best of 2 settings
52.0%
$0.15 / $0.60
154 Qwen3-MaxAlibaba
50.0%
$0.78 / $3.90
154 Llama 4 MaverickMeta
50.0%
$0.27 / $0.85
154 Step 3.5 FlashStepFun
50.0%
$0.10 / $0.30
157 Mistral Small 4Mistral AI · best of 2 settings
49.7%
$0.15 / $0.60
158 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
48.7%
$0.30 / $0.90
159 Qwen3-Coder-NextAlibaba
47.0%
$0.18 / $0.90
160 GLM-4.5-AirZ.ai
46.7%
$0.14 / $0.86
161 Qwen3-Coder-480B-A35BAlibaba
45.7%
$0.35 / $1.50
162 Granite 4.2 8BIBM
45.0%
$0.06 / $0.25
162 GPT-5 nano (high reasoning)OpenAI · best of 3 settings
45.0%
$0.05 / $0.40
164 GPT-4.1 miniOpenAI
44.0%
$0.40 / $1.60
165 Mercury 2Inception
43.7%
$0.25 / $0.75
166 Qwen3-Max-PreviewAlibaba
43.0%
–
166 o3-mini (high reasoning)OpenAI
43.0%
$1.10 / $4.40
168 Mistral Medium 3.1Mistral AI
42.7%
$0.40 / $2
169 GLM-4.7-FlashZ.ai · best of 2 settings
41.7%
$0.06 / $0.40
170 GPT-4o (2024-08-06)OpenAI
41.0%
$2.50 / $10
171 DeepSeek-V3-0324DeepSeek
40.7%
$0.25 / $1
172 Trinity Large ThinkingArcee AI
38.0%
$0.25 / $0.80
172 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA · best of 2 settings
38.0%
$0.05 / $0.20
174 Qwen3-4B-Thinking-2507Alibaba
37.3%
–
175 Mistral Large 3Mistral AI
36.0%
$0.50 / $1.50
176 gpt-oss-20b (high reasoning)OpenAI · best of 2 settings
34.7%
$0.03 / $0.15
177 Qwen3-235B-A22B-Instruct-2507Alibaba
33.9%
$0.15 / $0.75
178 Qwen3-VL-8B-ThinkingAlibaba
33.3%
$0.18 / $2.10
179 Qwen3-Coder-30B-A3B-InstructAlibaba
32.7%
$0.07 / $0.28
179 Qwen3-VL-235B-A22B-InstructAlibaba
32.7%
$0.30 / $1.50
181 Devstral 2Mistral AI
32.3%
$0.40 / $2
181 Solar Pro 3Upstage
32.3%
$0.15 / $0.60
183 Gemma 4 E4BGoogle · best of 2 settings
32.0%
–
183 Devstral MediumMistral AI
32.0%
–
185 Mistral Medium 3Mistral AI
31.3%
$0.40 / $2
185 Grok 4.1 Fast (non-reasoning, no reasoning)xAI
31.3%
–
187 DeepSeek-V3DeepSeek
29.3%
$0.26 / $1.03
188 Devstral Small 2Mistral AI
28.0%
–
188 Kimi Linear 48B A3B InstructMoonshot AI
28.0%
–
190 Llama 4 ScoutMeta
27.7%
$0.18 / $0.59
191 Qwen3-30B-A3B-Instruct-2507Alibaba
26.3%
$0.09 / $0.30
191 Ministral 3 14BMistral AI
26.3%
$0.20 / $0.20
193 Ministral 3 8BMistral AI
25.7%
$0.15 / $0.15
194 Grok 4.20 (0309, non-reasoning, no reasoning)xAI
25.3%
–
195 Grok 4 Fast (non-reasoning, no reasoning)xAI
24.0%
–
196 Mistral Small 3.1 24BMistral AI
22.3%
$0.35 / $0.56
197 Qwen3-VL-4B-ThinkingAlibaba
21.3%
–
197 Command ACohere
21.3%
$2.50 / $10
199 Qwen3.5-2B (reasoning on)Alibaba · best of 2 settings
21.0%
–
199 Nova Pro 1.0Amazon
21.0%
$0.80 / $3.20
201 Mistral Small 3.2 24BMistral AI
20.3%
$0.094 / $0.25
201 GPT-4.1 nanoOpenAI
20.3%
$0.10 / $0.40
203 DiffusionGemma 26B A4BGoogle
19.7%
–
204 Magistral Small 1.2Mistral AI
19.3%
–
205 Nova Lite 1.0Amazon
19.0%
$0.06 / $0.24
206 Devstral Small 1.1Mistral AI
18.7%
–
207 Llama 3.1 8B InstructMeta
18.0%
$0.05 / $0.08
208 Ministral 3 3BMistral AI
17.0%
$0.10 / $0.10
209 Qwen3-VL-8B-InstructAlibaba
16.7%
$0.12 / $0.46
210 Gemma 4 E2B (no reasoning)Google · best of 2 settings
16.3%
–
211 Llama 3.3 70B InstructMeta
15.7%
$0.59 / $0.79
212 Qwen3-VL-4B-InstructAlibaba
14.0%
–
213 Nova Micro 1.0Amazon
13.0%
$0.035 / $0.14
214 Qwen3-4B-Instruct-2507Alibaba
11.3%
–
215 Qwen3.5-0.8B (reasoning on)Alibaba · best of 2 settings
9.0%
–
216 Gemma 3 12BGoogle
8.3%
$0.05 / $0.15
217 Gemma 3 27BGoogle
7.3%
$0.12 / $0.20
218 Gemma 3 4BGoogle
6.7%
$0.05 / $0.10
218 Llama 3.2 1B InstructMeta
6.7%
$0.027 / $0.20
220 Llama 3.2 3B InstructMeta
4.3%
$0.05 / $0.33
221 Mistral Large 2 (2407)Mistral AI
2.0%
$2 / $6
222 Qwen3-14B (no reasoning)Alibaba · best of 2 settings
0.0%
$0.12 / $0.24
222 Qwen3-235B-A22B (no reasoning)Alibaba · best of 2 settings
0.0%
$0.46 / $1.82
222 Qwen3-30B-A3B (no reasoning)Alibaba · best of 2 settings
0.0%
$0.12 / $0.50
222 Qwen3-32B (no reasoning)Alibaba · best of 2 settings
0.0%
$0.14 / $0.40
222 Qwen3-8B (no reasoning)Alibaba · best of 2 settings
0.0%
$0.12 / $0.46
222 Qwen3-Omni-30B-A3B-InstructAlibaba
0.0%
–
222 Qwen3-Omni-30B-A3B-ThinkingAlibaba
0.0%
–
222 Gemma 3 270MGoogle
0.0%
–
222 Phi-4Microsoft
0.0%
$0.07 / $0.14
222 Mistral Small 3Mistral AI
0.0%
$0.05 / $0.08
222 Reka Flash 3Reka AI
0.0%
$0.10 / $0.20
222 GLM-4.5V (no reasoning)Z.ai · best of 2 settings
0.0%
$0.60 / $1.80

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of questions answered correctly over long documents, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not retrieval over your own document store.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.