Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

APEX-Agents-AA benchmark: model scores and methodology

Artificial Analysis's Stirrup-based implementation of long-horizon professional-services tasks across realistic workplace tools.

Current published leader
Gemini 3.5 Flash (high)
Top score
47%
Primary metric
Rubric success rate · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current APEX-Agents-AA model scores

15 configurations · 15 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 15 current configurations with a reported APEX-Agents-AA score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on APEX-Agents-AA

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. Gemini 3.5 Flash high 47%
  2. Kimi K3 41%
  3. GPT-5.6 Terra 39%
  4. GPT-5.6 Luna 36%
  5. GLM-5.2 max 34%
  6. Gemini 3.1 Pro (Preview) 32%
  7. DeepSeek V4 Pro Reasoning, Max Effort 24%
  8. Qwen3.7 Plus 22%
  9. Qwen3.5 397B A17B Reasoning 15%
  10. Step 3.7 Flash 15%
  11. Gemini 3.1 Flash-Lite 12%
  12. gpt-oss-120b high 3%
  13. MiMo-V2.5-Pro 2%
  14. NVIDIA Nemotron 3 Super 120B A12B Reasoning 2%
  15. gpt-oss-20b high 1%
0.0 23.5 47.1

Rubric success rate · higher is better

Current configurations with a reported APEX-Agents-AA value; missing values are omitted.

How to interpret the result

What do APEX-Agents-AA results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher rubric success rate is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read APEX-Agents-AA as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationRubric success rate
1 Gemini 3.5 Flash (high)Leader high 47%
2 Kimi K3 Published configuration 41%
3 GPT-5.6 Terra max 39%
4 GPT-5.6 Luna max 36%
5 GLM-5.2 (max) max 34%
6 Gemini 3.1 Pro (Preview) Published configuration 32%
7 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 24%
8 Qwen3.7 Plus Published configuration 22%
9 Qwen3.5 397B A17B (Reasoning) Reasoning 15%
10 Step 3.7 Flash Published configuration 15%
11 Gemini 3.1 Flash-Lite Published configuration 12%
12 gpt-oss-120b (high) high 3%
13 MiMo-V2.5-Pro Published configuration 2%
14 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 2%
15 gpt-oss-20b (high) high 1%
What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

Artificial Analysis's Stirrup-based implementation of long-horizon professional-services tasks across realistic workplace tools.

How to read the score

The published metric is Rubric success rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same apexAgents field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official APEX-Agents-AA resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗