Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

GPQA Diamond benchmark: model scores and methodology

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

Current published leader
GPT-5.6 Sol · max
Top score
94%5-way tie
Primary metric
Accuracy · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current GPQA Diamond model scores

250 configurations · 182 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 250 current configurations with a reported GPQA Diamond score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on GPQA Diamond

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. These results sit in a narrow band, so the scale below is zoomed — read each mark against the labelled axis, not the left edge. Marks are placed on the full source precision, so two models sharing a rounded label can still sit at different points. Score labels reproduce Artificial Analysis's rounded display value.

  1. GPT-5.6 Sol 94%
  2. Gemini 3.1 Pro (Preview) 94%
  3. Claude Opus 5 Adaptive Reasoning, High Effort 94%
  4. Kimi K3 94%
  5. Grok 4.5 93%
  6. MiniMax-M3 93%
  7. Gemini 3.6 Flash high 93%
  8. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 93%
  9. GPT-5.6 Terra 93%
  10. Qwen3.7 Max 92%
  11. Gemini 3.5 Flash high 92%
  12. GPT-5.3 Codex xhigh 92%
  13. Claude Sonnet 5 Adaptive Reasoning, Max Effort 91%
  14. GPT-5.6 Luna 91%
  15. DeepSeek V4 Pro Reasoning, High Effort 91%
  16. Qwen3.7 Plus 90%
  17. Muse Spark 1.1 xhigh 90%
  18. Hy3 90%
  19. Kimi K2.7 Code 90%
  20. GLM-5.2 max 89%
87.9 91.3 94.7

Accuracy · higher is better · zoomed scale

Current configurations with a reported GPQA Diamond value; missing values are omitted.

How to interpret the result

What do GPQA Diamond results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher accuracy is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read GPQA Diamond as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationAccuracy
1 GPT-5.6 SolLeader max 94%
2 Gemini 3.1 Pro (Preview) Published configuration 94%
3 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 94%
4 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 94%
5 Kimi K3 Published configuration 94%
6 Claude Opus 5 Adaptive Reasoning, Max Effort 93%
7 GPT-5.6 Sol (xhigh) xhigh 93%
8 Grok 4.5 high 93%
9 MiniMax-M3 Published configuration 93%
10 GPT-5.6 Sol (high) high 93%
11 Gemini 3.6 Flash (high) high 93%
12 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 93%
13 GPT-5.6 Sol (medium) medium 93%
14 GPT-5.6 Terra max 93%
15 Qwen3.7 Max Published configuration 92%
16 Gemini 3.5 Flash (high) high 92%
17 Gemini 3.5 Flash (medium) medium 92%
18 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 92%
19 GPT-5.3 Codex (xhigh) xhigh 92%
20 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 91%
21 GPT-5.6 Luna max 91%
22 GPT-5.6 Terra (xhigh) xhigh 91%
23 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 91%
24 Qwen3.7 Plus Published configuration 90%
25 GPT-5.6 Sol (low) low 90%
Show the remaining 225 configurations
#ModelConfigurationAccuracy
26 Muse Spark 1.1 (xhigh) xhigh 90%
27 Hy3 Published configuration 90%
28 GPT-5.6 Terra (high) high 90%
29 Kimi K2.7 Code Published configuration 90%
30 GLM-5.2 (max) max 89%
31 GPT-5.6 Luna (xhigh) xhigh 89%
32 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 89%
33 Qwen3.5 397B A17B (Reasoning) Reasoning 89%
34 GPT-5.6 Luna (high) high 89%
35 Nex-N2-Pro Published configuration 89%
36 Grok 4.3 (medium) medium 89%
37 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 89%
38 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 89%
39 Muse Spark Published configuration 88%
40 Qwen3.6 Plus Published configuration 88%
41 Agnes 2.5 Pro Alpha Published configuration 88%
42 GPT-5.6 Terra (medium) medium 87%
43 Inkling (xhigh) xhigh 87%
44 Motif 3 (Beta) Beta 87%
45 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 87%
46 Hy3-preview (Reasoning) Reasoning 87%
47 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 87%
48 MiMo-V2.5-Pro Published configuration 87%
49 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 86%
50 GPT-5.6 Luna (medium) medium 86%
51 Gemma 4 31B (Reasoning) Reasoning 86%
52 Qwen3.5 122B A10B (Reasoning) Reasoning 86%
53 Ring-2.6-1T Published configuration 86%
54 MiMo-V2-Omni-0327 Published configuration 85%
55 KAT Coder Pro V2 Published configuration 85%
56 MiMo-V2.5 Published configuration 85%
57 Nanbeige4.1-3B Published configuration 85%
58 JT-4.1 Flash 236B A21B Published configuration 85%
59 Gemini 2.5 Pro Published configuration 84%
60 GPT-5.6 Terra (low) low 84%
61 Grok 4.3 (low) low 84%
62 Qwen3.6 27B (Reasoning) Reasoning 84%
63 Qwen3.6 35B A3B (Reasoning) Reasoning 84%
64 Gemini 3.5 Flash-Lite Published configuration 84%
65 GPT-5.6 Luna (low) low 84%
66 MiMo-V2-Flash (Feb 2026) Feb 2026 84%
67 JT-35B-Flash Published configuration 83%
68 Qwen3.6 27B (Non-reasoning) Non-reasoning 83%
69 Gemini 3.5 Flash (minimal) minimal 83%
70 MiMo-V2-Omni Published configuration 83%
71 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 83%
72 o3 Published configuration 83%
73 Qwen3.5 Omni Plus Published configuration 83%
74 GPT-5.5 Instant (June 2026) June 2026 82%
75 Gemini 3.1 Flash-Lite Published configuration 82%
76 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 82%
77 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 82%
78 ERNIE 4.5 300B A47B Published configuration 81%
79 Nova 2.0 Lite (high) high 81%
80 Step 3.7 Flash Published configuration 81%
81 Qwen3.5 9B (Reasoning) Reasoning 81%
82 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 80%
83 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 80%
84 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 80%
85 EXAONE 4.5 33B Published configuration 79%
86 Gemma 4 26B A4B (Reasoning) Reasoning 79%
87 GPT-5.6 Sol (Non-reasoning) Non-reasoning 79%
88 Kimi K2.6 (Non-reasoning) Non-reasoning 79%
89 Qwen3.5 9B (Non-reasoning) Non-reasoning 79%
90 Nova 2.0 Pro Preview (medium) medium 78%
91 K-EXAONE (Reasoning) Reasoning 78%
92 gpt-oss-120b (high) high 78%
93 LongCat 2.0 Published configuration 78%
94 ERNIE 5.0 Thinking Preview Published configuration 78%
95 Mercury 2 Published configuration 77%
96 Mistral Small 4 (Reasoning) Reasoning 77%
97 Cogito v2.1 (Reasoning) Reasoning 77%
98 Nova 2.0 Lite (medium) medium 77%
99 Doubao Seed Code Published configuration 76%
100 KAT-Coder-Pro V1 Published configuration 76%
101 Gemma 4 31B (Non-reasoning) Non-reasoning 76%
102 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 76%
103 Command A+ Published configuration 76%
104 INTELLECT-3 Published configuration 76%
105 Nova 2.0 Omni (medium) medium 76%
106 Qwen3 Next 80B A3B (Reasoning) Reasoning 76%
107 Nemotron Cascade 2 30B A3B Published configuration 76%
108 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 76%
109 North Mini Code Published configuration 76%
110 Gemma 4 12B (Reasoning) Reasoning 75%
111 Ling-2.6-1T Published configuration 75%
112 Trinity Large Thinking Published configuration 75%
113 Nova 2.0 Pro Preview (low) low 75%
114 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 75%
115 Mistral Medium 3.5 Published configuration 75%
116 GPT-5.6 Terra (Non-reasoning) Non-reasoning 75%
117 Qwen3.5 Omni Flash Published configuration 74%
118 EXAONE 4.0 32B (Reasoning) Reasoning 74%
119 Magistral Medium 1.2 Published configuration 74%
120 Qwen3 Next 80B A3B Instruct Published configuration 74%
121 Sarvam 105B (high) high 74%
122 Qwen3 Coder Next Published configuration 74%
123 Apriel-v1.6-15B-Thinker Published configuration 73%
124 HyperNova 60B 2605 Published configuration 73%
125 Hy3-preview (Non-reasoning) Non-reasoning 73%
126 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 73%
127 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 73%
128 Qwen3 Omni 30B A3B (Reasoning) Reasoning 73%
129 Ring-flash-2.0 Published configuration 73%
130 Solar Pro 3 Published configuration 72%
131 Mi:dm K 2.5 Pro Preview Published configuration 72%
132 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 72%
133 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 72%
134 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 71%
135 K2 Think V2 Published configuration 71%
136 Qwen3.5 4B (Non-reasoning) Non-reasoning 71%
137 Mi:dm K 2.5 Pro Published configuration 70%
138 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 70%
139 Nova 2.0 Omni (low) low 70%
140 Nova 2.0 Lite (low) low 70%
141 K-EXAONE (Non-reasoning) Non-reasoning 69%
142 Motif-2-12.7B-Reasoning Published configuration 69%
143 Step3 VL 10B Published configuration 69%
144 gpt-oss-20b (high) high 69%
145 GLM-5.2 (Non-reasoning) Non-reasoning 69%
146 K2-V2 (high) high 68%
147 Mistral Large 3 Published configuration 68%
148 Qwen3.5 4B (Reasoning) Reasoning 68%
149 JT-MINI Published configuration 68%
150 Claude 4.5 Haiku (Reasoning) Reasoning 67%
151 gpt-oss-120b (low) low 67%
152 Llama 4 Maverick Published configuration 67%
153 DiffusionGemma 26B A4B Published configuration 67%
154 Magistral Small 1.2 Published configuration 66%
155 Falcon-H1R-7B Published configuration 66%
156 Gemma 4 12B (Non-reasoning) Non-reasoning 66%
157 Grok 4.3 (Non-reasoning) Non-reasoning 66%
158 Solar Open 100B (Reasoning) Reasoning 66%
159 MiMo-V2-Flash (Non-reasoning) Non-reasoning 66%
160 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 65%
161 GPT-5.6 Luna (Non-reasoning) Non-reasoning 65%
162 LongCat Flash Lite Published configuration 64%
163 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 64%
164 Sarvam 30B (high) high 63%
165 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 63%
166 Qwen3 Omni 30B A3B Instruct Published configuration 62%
167 HyperCLOVA X SEED Think (32B) 32B 62%
168 gpt-oss-20b (low) low 61%
169 Nova 2.0 Lite (Non-reasoning) Non-reasoning 60%
170 Tri-21B-Think Published configuration 60%
171 K2-V2 (medium) medium 60%
172 Devstral 2 Published configuration 59%
173 Ling 2.6 Flash Published configuration 59%
174 Olmo 3.1 32B Think Published configuration 59%
175 Llama 4 Scout Published configuration 59%
176 Phi-4 Published configuration 57%
177 Ministral 3 14B Published configuration 57%
178 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 57%
179 Mistral Small 4 (Non-reasoning) Non-reasoning 57%
180 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 57%
181 Nova Premier Published configuration 57%
182 Ling-mini-2.0 Published configuration 56%
183 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 56%
184 Nova 2.0 Omni (Non-reasoning) Non-reasoning 55%
185 Gemma 4 E4B (Non-reasoning) Non-reasoning 55%
186 K2-V2 (low) low 54%
187 Olmo 3.1 32B Instruct Published configuration 54%
188 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 54%
189 Devstral Small 2 Published configuration 53%
190 Reka Flash 3 Published configuration 53%
191 Command A Published configuration 53%
192 Gemma 4 E4B (Reasoning) Reasoning 52%
193 Olmo 3 7B Think Published configuration 52%
194 Exaone 4.0 1.2B (Reasoning) Reasoning 52%
195 Llama 3.1 Instruct 405B Published configuration 52%
196 NVIDIA Nemotron 3 Nano 4B Published configuration 51%
197 Llama 3.3 Instruct 70B Published configuration 50%
198 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 49%
199 Granite 4.1 30B Published configuration 48%
200 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 48%
201 LFM2 24B A2B Published configuration 47%
202 Ministral 3 8B Published configuration 47%
203 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 47%
204 LFM2.5-8B-A1B Published configuration 47%
205 Llama 3.1 Nemotron Instruct 70B Published configuration 46%
206 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 44%
207 G9v3-3B Published configuration 44%
208 Qwen3.5 2B (Non-reasoning) Non-reasoning 44%
209 Granite 4.1 8B Published configuration 43%
210 Llama 3.2 Instruct 90B (Vision) Vision 43%
211 Molmo2-8B Published configuration 43%
212 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 42%
213 Granite 4.0 H Small Published configuration 42%
214 Kimi Linear 48B A3B Instruct Published configuration 41%
215 Gemma 4 E2B (Non-reasoning) Non-reasoning 41%
216 Gemma 4 E2B (Reasoning) Reasoning 41%
217 Olmo 3 7B Instruct Published configuration 40%
218 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 40%
219 Jamba 1.7 Large Published configuration 39%
220 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) Non-reasoning 38%
221 Nova Micro Published configuration 36%
222 LFM2 8B A1B Published configuration 34%
223 LFM2.5-1.2B-Thinking Published configuration 34%
224 Jamba Reasoning 3B Published configuration 33%
225 Phi-4 Mini Instruct Published configuration 33%
226 Qwen3.5 2B (Reasoning) Reasoning 33%
227 Jamba 1.7 Mini Published configuration 32%
228 LFM2 2.6B Published configuration 32%
229 Phi-4 Multimodal Instruct Published configuration 32%
230 Granite 4.1 3B Published configuration 31%
231 MiniCPM-V 4.6 1.3B Published configuration 31%
232 Tiny Aya Global Published configuration 31%
233 Granite 4.0 Micro Published configuration 30%
234 Ministral 3 3B Published configuration 30%
235 LFM2.5-1.2B-Instruct Published configuration 29%
236 Granite 4.0 H 350M Published configuration 29%
237 LFM2.5-VL-1.6B Published configuration 29%
238 Granite 4.0 1B Published configuration 28%
239 MiniCPM5-1B (Reasoning) Reasoning 28%
240 Apertus 70B Instruct Published configuration 27%
241 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) Non-reasoning 27%
242 MiniCPM5-1B (Non-reasoning) Non-reasoning 27%
243 Apertus 8B Instruct Published configuration 26%
244 Granite 4.0 H 1B Published configuration 25%
245 Molmo 7B-D Published configuration 24%
246 Granite 4.0 350M Published configuration 24%
247 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 24%
248 Gemma 3 270M Published configuration 22%
249 Llama 3.2 Instruct 11B (Vision) Vision 22%
250 Qwen3.5 0.8B (Reasoning) Reasoning 12%

Showing the top 25 of 250 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

The most challenging subset of Graduate-Level Google-Proof Q&A, focused on scientific reasoning.

How to read the score

The published metric is Accuracy. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same gpqa field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official GPQA Diamond resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗