Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of graduate-level science questions answered correctly, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
359 configurations of 231 models
Unit
% of questions
Licence
Artificial Analysis commercial data licence

GPQA Diamond: Qwen3-Max-Thinking

Top 15 of 231 results · % of questions, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 GPT-6 Astra (extra-high reasoning)OpenAI 96.3%
  2. 2 Gemini 3.8 Flash (high reasoning)Google 95.3%
  3. 3 Grok 4.6 (high reasoning)xAI 94.9%
  4. 4 Gemini 3.7 Flash (high reasoning)Google 94.5%
  5. 5 Gemini 3.1 Pro PreviewGoogle 94.1%
  6. 5 Muse Spark 1.3 (extra-high reasoning)Meta 94.1%
  7. 5 GPT-5.6 Sol (max reasoning)OpenAI 94.1%
  8. 8 Claude Fable 5.1 (max reasoning)Anthropic 93.7%
  9. 8 Claude Opus 5 (high reasoning)Anthropic 93.7%
  10. 10 Qwen3.8-2.4T-A95BAlibaba 93.5%
  11. 10 Kimi K3 (max reasoning)Moonshot AI 93.5%
  12. 10 GPT-5.5 (extra-high reasoning)OpenAI 93.5%
  13. 13 Grok 4.5 (high reasoning)xAI 93.1%
  14. 14 MiniMax-M3MiniMax 92.9%
  15. 15 Qwen3.8-Max (0902)Alibaba 92.8%
  16. 71 Qwen3-Max-ThinkingAlibaba 86.1%

Full results

Artificial Analysis: GPQA Diamond, % of questions, higher is better
#ModelGPQA Diamond
% of questions, higher is better
Price
$ per million tokens, in / out
1 GPT-6 Astra (extra-high reasoning)OpenAI · best of 5 settings
96.3%
$10 / $50
2 Gemini 3.8 Flash (high reasoning)Google · best of 3 settings
95.3%
$1.50 / $7.50
3 Grok 4.6 (high reasoning)xAI · best of 4 settings
94.9%
$2 / $6
4 Gemini 3.7 Flash (high reasoning)Google · best of 3 settings
94.5%
$1.50 / $7.50
5 Gemini 3.1 Pro PreviewGoogle
94.1%
$2 / $12
5 Muse Spark 1.3 (extra-high reasoning)Meta · best of 2 settings
94.1%
$1.25 / $4.25
5 GPT-5.6 Sol (max reasoning)OpenAI · best of 6 settings
94.1%
$4 / $20
8 Claude Fable 5.1 (max reasoning)Anthropic · best of 5 settings
93.7%
$10 / $50
8 Claude Opus 5 (high reasoning)Anthropic · best of 5 settings
93.7%
$5 / $25
10 Qwen3.8-2.4T-A95BAlibaba
93.5%
$2 / $6
10 Kimi K3 (max reasoning)Moonshot AI · best of 2 settings
93.5%
$3 / $15
10 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
93.5%
$5 / $30
13 Grok 4.5 (high reasoning)xAI
93.1%
$2 / $6
14 MiniMax-M3MiniMax
92.9%
$0.30 / $1.20
15 Qwen3.8-Max (0902)Alibaba
92.8%
$2 / $6
15 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek
92.8%
$1.32 / $3.96
15 Gemini 3.6 Flash (high reasoning)Google
92.8%
$1.50 / $7.50
18 Qwen3.8-Max (0803)Alibaba
92.7%
–
19 Claude Fable 5 (max reasoning)Anthropic
92.6%
$10 / $50
20 GPT-5.6 Terra (max reasoning)OpenAI · best of 6 settings
92.5%
$2 / $12
21 Qwen3.7-MaxAlibaba
92.3%
$1.48 / $4.42
21 Qwen3.8-Flash-NextAlibaba
92.3%
–
23 Gemini 3.5 Flash (high reasoning)Google · best of 3 settings
92.2%
$1.50 / $9
24 Claude Opus 4.8 (max reasoning)Anthropic
92.0%
$5 / $25
24 GPT-5.4 (extra-high reasoning)OpenAI · best of 3 settings
92.0%
$2.50 / $15
26 GLM-5.3 (max reasoning)Z.ai
91.7%
$1.40 / $4.40
27 GPT-5.3-Codex (extra-high reasoning)OpenAI
91.5%
$1.75 / $14
28 Claude Opus 4.7 (max reasoning)Anthropic · best of 2 settings
91.4%
$5 / $25
29 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
91.3%
$0.44 / $1.32
30 GLM-5.3-FlashZ.ai
91.2%
$0.15 / $0.50
31 Claude Sonnet 5 (max reasoning)Anthropic · best of 2 settings
91.1%
$2 / $10
31 Kimi K2.6Moonshot AI · best of 2 settings
91.1%
$0.95 / $4
31 GPT-5.6 Luna (max reasoning)OpenAI · best of 6 settings
91.1%
$0.20 / $1.20
31 Grok 4.20xAI · best of 2 settings
91.1%
$1.25 / $2.50
35 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
90.8%
$0.14 / $0.28
35 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
90.8%
–
37 Qwen3.8-27B (extra-high reasoning)Alibaba · best of 4 settings
90.5%
$0.50 / $3
37 DeepSeek-V4-Pro (0423, high reasoning)DeepSeek · best of 3 settings
90.5%
$1.42 / $2.83
39 Muse Spark 1.2 (extra-high reasoning)Meta
90.4%
$1.25 / $4.25
40 GPT-5.2 (extra-high reasoning)OpenAI · best of 3 settings
90.3%
$1.75 / $14
41 Grok 4.3 (high reasoning)xAI · best of 4 settings
90.1%
$1.25 / $2.50
42 Qwen3.7-PlusAlibaba
90.0%
$0.32 / $1.28
43 GPT-5.2-Codex (extra-high reasoning)OpenAI
89.9%
$1.75 / $14
44 Gemini 3 Flash Preview (reasoning on)Google · best of 2 settings
89.8%
$0.50 / $3
44 Muse Spark 1.1 (extra-high reasoning)Meta
89.8%
$1.25 / $4.25
46 Hy3Tencent
89.7%
$0.14 / $0.58
47 Claude Opus 4.6 (max reasoning)Anthropic · best of 2 settings
89.6%
$5 / $25
47 Kimi K2.7 CodeMoonshot AI
89.6%
$0.95 / $4
49 Inkling SmallThinking Machines
89.5%
$0.45 / $1.20
49 Grok Build 0.1xAI
89.5%
$1 / $2
49 GLM-5.2 (max reasoning)Z.ai · best of 2 settings
89.5%
$1.40 / $4.40
52 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek · best of 3 settings
89.4%
$0.14 / $0.28
53 Qwen3.5-397B-A17BAlibaba · best of 2 settings
89.3%
$0.55 / $3.50
54 Solar Pro 4Upstage
89.1%
$0.09 / $0.36
55 Qwen3.6-Max-PreviewAlibaba
88.8%
$1.03 / $6.16
56 Grok 4.20 (0309, reasoning)xAI
88.5%
–
57 Muse SparkMeta
88.4%
–
58 Qwen3.6-PlusAlibaba
88.2%
$0.33 / $1.95
59 Kimi K2.5Moonshot AI · best of 2 settings
87.9%
$0.57 / $2.85
60 Grok 4xAI
87.7%
–
61 Claude Sonnet 4.6 (max reasoning)Anthropic · best of 2 settings
87.5%
$3 / $15
61 GPT-5.4 mini (extra-high reasoning)OpenAI · best of 3 settings
87.5%
$0.75 / $4.50
63 MiniMax-M2.7MiniMax
87.4%
$0.30 / $1.20
64 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
87.3%
$1.25 / $10
65 Inkling (extra-high reasoning)Thinking Machines
87.2%
$0.95 / $4.05
66 DeepSeek-V3.2-SpecialeDeepSeek
87.1%
–
67 GLM-5.1Z.ai · best of 2 settings
86.8%
$1.38 / $4.40
68 Hy3 Preview (reasoning on)Tencent · best of 2 settings
86.7%
$0.18 / $0.60
69 Claude Opus 4.5 (reasoning on)Anthropic · best of 2 settings
86.6%
$5 / $25
69 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
86.6%
$0.43 / $0.87
71 Qwen3-Max-ThinkingAlibaba
86.1%
$0.78 / $3.90
72 GPT-5.1-Codex (high reasoning)OpenAI
86.0%
$1.25 / $10
73 GLM-4.7Z.ai · best of 2 settings
85.9%
$0.54 / $1.98
74 Qwen3.5-27BAlibaba · best of 2 settings
85.8%
$0.27 / $2.16
75 Qwen3.5-122B-A10BAlibaba · best of 2 settings
85.7%
$0.26 / $2.08
75 Gemma 4 31BGoogle · best of 2 settings
85.7%
$0.14 / $0.40
77 GPT-5 (high reasoning)OpenAI · best of 4 settings
85.4%
$1.25 / $10
78 Grok 4.1 Fast (reasoning)xAI
85.3%
–
79 MiMo-V2.5Xiaomi
84.9%
$0.17 / $0.34
80 MiniMax-M2.5MiniMax
84.8%
$0.30 / $1.20
81 Grok 4 Fast (reasoning)xAI
84.7%
–
81 GLM-5-TurboZ.ai
84.7%
$1.20 / $4
83 GPT-5.5 Instant (2026-05-26)OpenAI · best of 2 settings
84.6%
–
84 Qwen3.5-35B-A3BAlibaba · best of 2 settings
84.5%
$0.16 / $1.30
84 o3-proOpenAI
84.5%
$20 / $80
86 Gemini 2.5 ProGoogle
84.4%
$1.25 / $10
87 Qwen3.6-27BAlibaba · best of 2 settings
84.2%
$0.30 / $3.20
88 Qwen3.6-35B-A3BAlibaba · best of 2 settings
84.1%
$0.10 / $1
89 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
84.0%
$0.30 / $0.96
90 Gemini 3.5 Flash-LiteGoogle
83.8%
$0.30 / $2.50
90 Kimi K2 ThinkingMoonshot AI
83.8%
$0.60 / $2.50
92 GPT-5-Codex (high reasoning)OpenAI
83.7%
–
93 Muse Glimmer (high reasoning)Meta
83.5%
–
94 Claude Sonnet 4.5 (reasoning on)Anthropic · best of 2 settings
83.4%
$3 / $15
95 Step 3.5 FlashStepFun
83.1%
$0.10 / $0.30
96 MiniMax-M2.1MiniMax
83.0%
$0.30 / $1.20
97 GPT-5 mini (high reasoning)OpenAI · best of 3 settings
82.8%
$0.25 / $2
98 o3OpenAI
82.7%
$2 / $8
99 Qwen3.5-Omni-PlusAlibaba
82.6%
–
100 Gemini 3.1 Flash-Lite PreviewGoogle
82.2%
$0.25 / $1.50
101 GLM-5Z.ai · best of 2 settings
82.0%
$0.95 / $2.55
102 GPT-5.4 nano (extra-high reasoning)OpenAI · best of 3 settings
81.7%
$0.20 / $1.25
103 DeepSeek-R1-0528DeepSeek
81.3%
$0.50 / $2.18
103 GPT-5.1-Codex-Mini (high reasoning)OpenAI
81.3%
$0.25 / $2
105 Nova 2 Lite (high reasoning)Amazon · best of 4 settings
81.1%
$0.30 / $2.50
106 Claude Opus 4.1 (reasoning on)Anthropic
80.9%
$15 / $75
106 GLM-5V-TurboZ.ai
80.9%
$1.20 / $4
108 Qwen3.5-9BAlibaba · best of 2 settings
80.6%
$0.10 / $0.15
109 Nemotron 3 Super 120B A12BNVIDIA
80.0%
$0.085 / $0.40
110 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
79.7%
$0.27 / $0.41
111 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google · best of 2 settings
79.3%
–
112 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek · best of 2 settings
79.2%
$0.27 / $1
112 Gemma 4 26B A4BGoogle · best of 2 settings
79.2%
$0.10 / $0.30
114 Qwen3-235B-A22B-Thinking-2507Alibaba
79.0%
$0.30 / $3
114 Gemini 2.5 Flash (reasoning on)Google · best of 2 settings
79.0%
$0.30 / $2.50
116 Grok 4.20 (0309, non-reasoning, no reasoning)xAI
78.5%
–
117 o4-mini (high reasoning)OpenAI
78.4%
$1.10 / $4.40
118 gpt-oss-120b (high reasoning)OpenAI · best of 2 settings
78.2%
$0.15 / $0.60
118 GLM-4.5Z.ai
78.2%
$0.60 / $2.20
120 GLM-4.6 (reasoning on)Z.ai · best of 2 settings
78.0%
$0.50 / $2
121 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
77.9%
$0.55 / $1.65
122 Claude Sonnet 4 (reasoning on)Anthropic · best of 2 settings
77.7%
$3 / $15
122 MiniMax-M2MiniMax
77.7%
$0.30 / $1.20
124 Qwen3-Max-Thinking-PreviewAlibaba
77.6%
–
125 o3-mini (high reasoning)OpenAI · best of 2 settings
77.3%
$1.10 / $4.40
126 Qwen3-VL-235B-A22B-ThinkingAlibaba
77.2%
$0.40 / $4
127 Qwen3.5-4B (reasoning on)Alibaba · best of 2 settings
77.1%
–
128 Mercury 2Inception
77.0%
$0.25 / $0.75
129 Mistral Small 4Mistral AI · best of 2 settings
76.9%
$0.15 / $0.60
130 Kimi K2 (0905)Moonshot AI
76.7%
$0.60 / $2.50
131 Kimi K2 (0711)Moonshot AI
76.6%
$0.57 / $2.30
132 Qwen3-Max-PreviewAlibaba
76.4%
–
132 Qwen3-MaxAlibaba
76.4%
$0.78 / $3.90
134 Qwen3-Next-80B-A3B-ThinkingAlibaba
75.9%
$0.15 / $1.20
135 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA · best of 2 settings
75.7%
$0.05 / $0.20
136 Qwen3-235B-A22B-Instruct-2507Alibaba
75.3%
$0.15 / $0.75
136 Gemma 4 12BGoogle · best of 2 settings
75.3%
–
138 Trinity Large ThinkingArcee AI
75.2%
$0.25 / $0.80
139 Mistral Medium 3.5Mistral AI
74.8%
$1.50 / $7.50
140 o1OpenAI
74.7%
$15 / $60
141 Qwen3.5-Omni-FlashAlibaba
74.2%
–
142 Magistral Medium 1.2Mistral AI
73.9%
–
143 Qwen3-Next-80B-A3B-InstructAlibaba
73.8%
$0.10 / $1.10
144 Qwen3-Coder-NextAlibaba
73.7%
$0.18 / $0.90
145 Qwen3-VL-32B-ThinkingAlibaba
73.3%
–
145 GLM-4.5-AirZ.ai
73.3%
$0.14 / $0.86
147 Grok Code Fast 1xAI
72.7%
–
148 Qwen3-Omni-30B-A3B-ThinkingAlibaba
72.6%
–
149 Solar Pro 3Upstage
72.4%
$0.15 / $0.60
150 Qwen3-VL-30B-A3B-ThinkingAlibaba
72.0%
$0.29 / $1
151 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
71.9%
$0.30 / $0.90
152 Qwen3-VL-235B-A22B-InstructAlibaba
71.2%
$0.30 / $1.50
153 Gemini 2.5 Flash-Lite Preview (09-2025, reasoning on)Google · best of 2 settings
70.9%
–
154 DeepSeek-R1DeepSeek
70.8%
$0.70 / $2.50
155 Qwen3-30B-A3B-Thinking-2507Alibaba
70.7%
$0.20 / $2.40
156 Qwen3-235B-A22B (reasoning on)Alibaba · best of 2 settings
70.0%
$0.46 / $1.82
157 Qwen3-VL-30B-A3B-InstructAlibaba
69.5%
$0.15 / $0.60
158 gpt-oss-20b (high reasoning)OpenAI · best of 2 settings
68.8%
$0.03 / $0.15
159 GPT-5 ChatOpenAI
68.6%
–
160 GLM-4.5V (reasoning on)Z.ai · best of 2 settings
68.4%
$0.60 / $1.80
161 Mistral Large 3Mistral AI
68.0%
$0.50 / $1.50
162 GPT-5 nano (high reasoning)OpenAI · best of 3 settings
67.6%
$0.05 / $0.40
163 Claude Haiku 4.5 (reasoning on)Anthropic · best of 2 settings
67.2%
$1 / $5
164 Qwen3-VL-32B-InstructAlibaba
67.1%
$0.10 / $0.42
164 Llama 4 MaverickMeta
67.1%
$0.27 / $0.85
166 DiffusionGemma 26B A4BGoogle
66.9%
–
167 Qwen3-32B (reasoning on)Alibaba · best of 2 settings
66.8%
$0.14 / $0.40
168 Qwen3-4B-Thinking-2507Alibaba
66.7%
–
169 GPT-4.1OpenAI
66.6%
$2 / $8
170 GPT-4.1 miniOpenAI
66.4%
$0.40 / $1.60
171 Magistral Small 1.2Mistral AI
66.3%
–
172 Qwen3-30B-A3B-Instruct-2507Alibaba
65.9%
$0.09 / $0.30
173 DeepSeek-V3-0324DeepSeek
65.5%
$0.25 / $1
174 Grok 4.1 Fast (non-reasoning, no reasoning)xAI
63.7%
–
175 Granite 4.2 8BIBM
63.1%
$0.06 / $0.25
176 Gemini 2.5 Flash-Lite (reasoning on)Google · best of 2 settings
62.5%
$0.10 / $0.40
177 Qwen3-Omni-30B-A3B-InstructAlibaba
62.0%
–
178 Qwen3-Coder-480B-A35BAlibaba
61.8%
$0.35 / $1.50
179 Qwen3-30B-A3B (reasoning on)Alibaba · best of 2 settings
61.6%
$0.12 / $0.50
180 Grok 4 Fast (non-reasoning, no reasoning)xAI
60.6%
–
181 Qwen3-14B (reasoning on)Alibaba · best of 2 settings
60.4%
$0.12 / $0.24
182 Devstral 2Mistral AI
59.4%
$0.40 / $2
183 Qwen3-8B (reasoning on)Alibaba · best of 2 settings
58.9%
$0.12 / $0.46
184 Mistral Medium 3.1Mistral AI
58.8%
$0.40 / $2
185 Llama 4 ScoutMeta
58.7%
$0.18 / $0.59
186 GLM-4.7-FlashZ.ai · best of 2 settings
58.1%
$0.06 / $0.40
187 Qwen3-VL-8B-ThinkingAlibaba
57.9%
$0.18 / $2.10
188 Mistral Medium 3Mistral AI
57.8%
$0.40 / $2
189 Gemma 4 E4BGoogle · best of 2 settings
57.6%
–
190 Phi-4Microsoft
57.5%
$0.07 / $0.14
191 Ministral 3 14BMistral AI
57.2%
$0.20 / $0.20
192 DeepSeek-V3DeepSeek
55.7%
$0.26 / $1.03
193 Devstral Small 2Mistral AI
53.2%
–
194 Reka Flash 3Reka AI
52.9%
$0.10 / $0.20
195 Command ACohere
52.7%
$2.50 / $10
196 GPT-4o (2024-05-13)OpenAI
52.6%
$5 / $15
197 GPT-4o (2024-08-06)OpenAI
52.1%
$2.50 / $10
198 Qwen3-4B-Instruct-2507Alibaba
51.7%
–
199 Qwen3-Coder-30B-A3B-InstructAlibaba
51.6%
$0.07 / $0.28
200 GPT-4.1 nanoOpenAI
51.2%
$0.10 / $0.40
201 Mistral Small 3.2 24BMistral AI
50.5%
$0.094 / $0.25
202 Nova Pro 1.0Amazon
49.9%
$0.80 / $3.20
203 Llama 3.3 70B InstructMeta
49.8%
$0.59 / $0.79
204 Qwen3-VL-4B-ThinkingAlibaba
49.4%
–
205 Devstral MediumMistral AI
49.2%
–
206 Qwen2.5-72B-InstructAlibaba
49.1%
$0.36 / $0.40
207 Mistral Large 2 (2407)Mistral AI
47.2%
$2 / $6
208 Ministral 3 8BMistral AI
47.1%
$0.15 / $0.15
209 Mistral Small 3Mistral AI
46.2%
$0.05 / $0.08
210 Qwen3.5-2B (reasoning on)Alibaba · best of 2 settings
45.6%
–
211 Mistral Small 3.1 24BMistral AI
45.4%
$0.35 / $0.56
212 Nova Lite 1.0Amazon
43.3%
$0.06 / $0.24
212 Gemma 4 E2BGoogle · best of 2 settings
43.3%
–
214 Gemma 3 27BGoogle
42.8%
$0.12 / $0.20
215 Qwen3-VL-8B-InstructAlibaba
42.7%
$0.12 / $0.46
216 GPT-4o mini (2024-07-18)OpenAI
42.6%
$0.15 / $0.60
217 Qwen2.5-Coder-32B-InstructAlibaba
41.7%
$0.66 / $1
218 Devstral Small 1.1Mistral AI
41.4%
–
219 Kimi Linear 48B A3B InstructMoonshot AI
41.2%
–
220 Llama 3.1 70B InstructMeta
40.9%
$0.40 / $0.40
221 Qwen3-VL-4B-InstructAlibaba
37.1%
–
222 Nova Micro 1.0Amazon
35.8%
$0.035 / $0.14
222 Ministral 3 3BMistral AI
35.8%
$0.10 / $0.10
224 Gemma 3 12BGoogle
34.9%
$0.05 / $0.15
225 Mixtral 8x22B InstructMistral AI
33.2%
$2 / $6
226 Gemma 3 4BGoogle
29.1%
$0.05 / $0.10
227 Llama 3.1 8B InstructMeta
25.9%
$0.05 / $0.08
228 Llama 3.2 3B InstructMeta
25.5%
$0.05 / $0.33
229 Qwen3.5-0.8B (no reasoning)Alibaba · best of 2 settings
23.6%
–
230 Gemma 3 270MGoogle
22.4%
–
231 Llama 3.2 1B InstructMeta
19.6%
$0.027 / $0.20

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of graduate-level science questions answered correctly, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not applied work; multiple-choice science questions.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.