Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

ITBench-AA benchmark: model scores and methodology

Artificial Analysis's implementation of Kubernetes incident root-cause analysis from offline SRE snapshots.

Current published leader
GPT-5.6 Sol · max
Top score
56%
Primary metric
Root-cause diagnosis score · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current ITBench-AA model scores

19 configurations · 19 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 19 current configurations with a reported ITBench-AA score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on ITBench-AA

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. GPT-5.6 Sol 56%
  2. GPT-5.6 Terra 51%
  3. Kimi K3 48%
  4. GLM-5.2 max 43%
  5. Qwen3.7 Max 42%
  6. Gemini 3.5 Flash high 40%
  7. GPT-5.6 Luna 40%
  8. DeepSeek V4 Pro Reasoning, Max Effort 38%
  9. MiMo-V2.5-Pro 38%
  10. Gemma 4 31B Reasoning 37%
  11. Qwen3.5 397B A17B Reasoning 34%
  12. DeepSeek V4 Flash Reasoning, Max Effort 32%
  13. Gemini 3.1 Pro (Preview) 30%
  14. Step 3.7 Flash 30%
  15. Claude 4.5 Haiku Reasoning 27%
  16. Gemma 4 26B A4B Reasoning 24%
  17. gpt-oss-120b high 6%
  18. NVIDIA Nemotron 3 Super 120B A12B Reasoning 1%
  19. Llama 3.3 Instruct 70B 1%
0.0 28.1 56.2

Root-cause diagnosis score · higher is better

Current configurations with a reported ITBench-AA value; missing values are omitted.

How to interpret the result

What do ITBench-AA results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher root-cause diagnosis score is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read ITBench-AA as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationRoot-cause diagnosis score
1 GPT-5.6 SolLeader max 56%
2 GPT-5.6 Terra max 51%
3 Kimi K3 Published configuration 48%
4 GLM-5.2 (max) max 43%
5 Qwen3.7 Max Published configuration 42%
6 Gemini 3.5 Flash (high) high 40%
7 GPT-5.6 Luna max 40%
8 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 38%
9 MiMo-V2.5-Pro Published configuration 38%
10 Gemma 4 31B (Reasoning) Reasoning 37%
11 Qwen3.5 397B A17B (Reasoning) Reasoning 34%
12 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 32%
13 Gemini 3.1 Pro (Preview) Published configuration 30%
14 Step 3.7 Flash Published configuration 30%
15 Claude 4.5 Haiku (Reasoning) Reasoning 27%
16 Gemma 4 26B A4B (Reasoning) Reasoning 24%
17 gpt-oss-120b (high) high 6%
18 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 1%
19 Llama 3.3 Instruct 70B Published configuration 1%
What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

Artificial Analysis's implementation of Kubernetes incident root-cause analysis from offline SRE snapshots.

How to read the score

The published metric is Root-cause diagnosis score. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same itbenchSre field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official ITBench-AA resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗