Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

GDPval-AA v2 benchmark: model scores and methodology

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

Current published leader
Claude Opus 5 · Adaptive Reasoning, Max Effort
Top score
68%
Primary metric
Normalized GDPval-AA v2 score · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current GDPval-AA v2 model scores

133 configurations · 96 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 133 current configurations with a reported GDPval-AA v2 score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on GDPval-AA v2

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. Claude Opus 5 68%
  2. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 62%
  3. GPT-5.6 Sol 62%
  4. Kimi K3 59%
  5. Claude Sonnet 5 Adaptive Reasoning, Max Effort 55%
  6. GPT-5.6 Terra 54%
  7. GPT-5.6 Luna 54%
  8. Grok 4.5 51%
  9. GLM-5.2 max 50%
  10. Gemini 3.6 Flash high 46%
  11. MiniMax-M3 45%
  12. Muse Spark 1.1 xhigh 44%
  13. Gemini 3.5 Flash high 42%
  14. DeepSeek V4 Pro Reasoning, Max Effort 40%
  15. Qwen3.7 Max 39%
  16. JT-4.1 Flash 236B A21B 38%
  17. MiMo-V2.5-Pro 38%
  18. Motif 3 Beta 38%
  19. Nex-N2-Pro 38%
  20. Inkling xhigh 37%
0.0 34.0 68.1

Normalized GDPval-AA v2 score · higher is better

Current configurations with a reported GDPval-AA v2 value; missing values are omitted.

How to interpret the result

What do GDPval-AA v2 results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher normalized gdpval-aa v2 score is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read GDPval-AA v2 as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationNormalized GDPval-AA v2 score
1 Claude Opus 5Leader Adaptive Reasoning, Max Effort 68%
2 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 66%
3 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 62%
4 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 62%
5 GPT-5.6 Sol max 62%
6 GPT-5.6 Sol (xhigh) xhigh 59%
7 Kimi K3 Published configuration 59%
8 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 57%
9 GPT-5.6 Sol (high) high 56%
10 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 55%
11 GPT-5.6 Terra max 54%
12 GPT-5.6 Luna max 54%
13 GPT-5.6 Terra (xhigh) xhigh 54%
14 GPT-5.6 Sol (medium) medium 53%
15 GPT-5.6 Luna (xhigh) xhigh 51%
16 Grok 4.5 high 51%
17 GLM-5.2 (max) max 50%
18 GPT-5.6 Terra (high) high 50%
19 Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 50%
20 GPT-5.6 Luna (high) high 48%
21 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 48%
22 GPT-5.6 Sol (low) low 47%
23 Gemini 3.6 Flash (high) high 46%
24 GPT-5.6 Terra (medium) medium 45%
25 Claude Sonnet 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 45%
Show the remaining 108 configurations
#ModelConfigurationNormalized GDPval-AA v2 score
26 MiniMax-M3 Published configuration 45%
27 GLM-5.2 (Non-reasoning) Non-reasoning 44%
28 GPT-5.6 Sol (Non-reasoning) Non-reasoning 44%
29 Muse Spark 1.1 (xhigh) xhigh 44%
30 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 43%
31 Gemini 3.5 Flash (high) high 42%
32 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 40%
33 Claude Sonnet 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 40%
34 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 40%
35 GPT-5.6 Luna (medium) medium 39%
36 Qwen3.7 Max Published configuration 39%
37 JT-4.1 Flash 236B A21B Published configuration 38%
38 MiMo-V2.5-Pro Published configuration 38%
39 Motif 3 (Beta) Beta 38%
40 GPT-5.6 Terra (low) low 38%
41 Nex-N2-Pro Published configuration 38%
42 GPT-5.6 Terra (Non-reasoning) Non-reasoning 37%
43 Inkling (xhigh) xhigh 37%
44 Hy3 Published configuration 36%
45 Claude Sonnet 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 36%
46 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 34%
47 Kimi K2.7 Code Published configuration 34%
48 Agnes 2.5 Pro Alpha Published configuration 33%
49 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 33%
50 GPT-5.6 Luna (low) low 33%
51 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 32%
52 MiMo-V2.5 Published configuration 32%
53 Muse Spark Published configuration 32%
54 Gemini 3.5 Flash-Lite Published configuration 32%
55 Qwen3.6 Plus Published configuration 32%
56 Qwen3.6 27B (Reasoning) Reasoning 32%
57 Qwen3.6 27B (Non-reasoning) Non-reasoning 30%
58 Grok 4.3 (Non-reasoning) Non-reasoning 30%
59 GPT-5.6 Luna (Non-reasoning) Non-reasoning 29%
60 Qwen3.6 35B A3B (Reasoning) Reasoning 28%
61 LongCat 2.0 Published configuration 26%
62 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 26%
63 Step 3.7 Flash Published configuration 26%
64 Qwen3.5 122B A10B (Reasoning) Reasoning 24%
65 Gemini 3.1 Pro (Preview) Published configuration 23%
66 Qwen3.5 397B A17B (Reasoning) Reasoning 23%
67 Qwen3.7 Plus Published configuration 22%
68 Mistral Medium 3.5 Published configuration 22%
69 Ring-2.6-1T Published configuration 21%
70 Claude 4.5 Haiku (Reasoning) Reasoning 21%
71 KAT Coder Pro V2 Published configuration 20%
72 KAT-Coder-Pro V1 Published configuration 20%
73 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 19%
74 G9v3-3B Published configuration 18%
75 MiMo-V2-Flash (Non-reasoning) Non-reasoning 17%
76 Gemma 4 31B (Reasoning) Reasoning 16%
77 gpt-oss-120b (high) high 15%
78 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 15%
79 Gemma 4 26B A4B (Reasoning) Reasoning 14%
80 Gemma 4 31B (Non-reasoning) Non-reasoning 13%
81 Devstral 2 Published configuration 12%
82 Devstral Small 2 Published configuration 12%
83 GPT-5.5 Instant (June 2026) June 2026 11%
84 Qwen3 Coder Next Published configuration 11%
85 Command A+ Published configuration 11%
86 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 10%
87 Mercury 2 Published configuration 10%
88 Nova 2.0 Pro Preview (medium) medium 9%
89 EXAONE 4.5 33B Published configuration 9%
90 Gemini 2.5 Pro Published configuration 9%
91 HyperNova 60B 2605 Published configuration 8%
92 Nova 2.0 Pro Preview (low) low 8%
93 Gemma 4 12B (Reasoning) Reasoning 8%
94 Qwen3.5 9B (Reasoning) Reasoning 7%
95 Gemini 3.1 Flash-Lite Published configuration 7%
96 Mistral Large 3 Published configuration 7%
97 K-EXAONE (Reasoning) Reasoning 5%
98 Nova 2.0 Lite (high) high 5%
99 Mistral Small 4 (Reasoning) Reasoning 5%
100 gpt-oss-20b (high) high 3%
101 Trinity Large Thinking Published configuration 3%
102 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 3%
103 DiffusionGemma 26B A4B Published configuration 3%
104 Ling 2.6 Flash Published configuration 3%
105 North Mini Code Published configuration 2%
106 Nemotron Cascade 2 30B A3B Published configuration 0%
107 Solar Pro 3 Published configuration 0%
108 Gemma 4 E2B (Reasoning) Reasoning 0%
109 Gemma 4 E4B (Reasoning) Reasoning 0%
110 Granite 4.1 30B Published configuration 0%
111 Granite 4.1 3B Published configuration 0%
112 K2 Think V2 Published configuration 0%
113 Llama 3.3 Instruct 70B Published configuration 0%
114 Llama 4 Maverick Published configuration 0%
115 Llama 4 Scout Published configuration 0%
116 Magistral Medium 1.2 Published configuration 0%
117 Magistral Small 1.2 Published configuration 0%
118 MiniCPM-V 4.6 1.3B Published configuration 0%
119 Ministral 3 14B Published configuration 0%
120 Ministral 3 3B Published configuration 0%
121 Ministral 3 8B Published configuration 0%
122 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 0%
123 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 0%
124 NVIDIA Nemotron 3 Nano 4B Published configuration 0%
125 Nanbeige4.1-3B Published configuration 0%
126 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 0%
127 Phi-4 Mini Instruct Published configuration 0%
128 Qwen3 Next 80B A3B (Reasoning) Reasoning 0%
129 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 0%
130 Qwen3.5 0.8B (Reasoning) Reasoning 0%
131 Qwen3.5 2B (Non-reasoning) Non-reasoning 0%
132 Qwen3.5 2B (Reasoning) Reasoning 0%
133 gpt-oss-120b (low) low 0%

Showing the top 25 of 133 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

Artificial Analysis's agentic evaluation of economically valuable, real-world work tasks based on the GDPval dataset.

How to read the score

The published metric is Normalized GDPval-AA v2 score. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same gdpvalNormalized field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official GDPval-AA v2 resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗