Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

Terminal-Bench v2.1 benchmark: model scores and methodology

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

Current published leader
GPT-5.6 Sol (xhigh)
Top score
90%
Primary metric
Pass rate · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current Terminal-Bench v2.1 model scores

131 configurations · 97 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 131 current configurations with a reported Terminal-Bench v2.1 score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on Terminal-Bench v2.1

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. GPT-5.6 Sol xhigh 90%
  2. Claude Opus 5 89%
  3. GPT-5.6 Terra 88%
  4. Kimi K3 85%
  5. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 85%
  6. Grok 4.5 82%
  7. GPT-5.6 Luna 81%
  8. Claude Sonnet 5 Adaptive Reasoning, Max Effort 81%
  9. Gemini 3.5 Flash high 79%
  10. GLM-5.2 max 78%
  11. Muse Spark 1.1 xhigh 78%
  12. Gemini 3.6 Flash high 78%
  13. Qwen3.7 Max 75%
  14. Gemini 3.1 Pro (Preview) 74%
  15. Motif 3 Beta 71%
  16. KAT Coder Pro V2 70%
  17. Nex-N2-Pro 68%
  18. Kimi K2.7 Code 67%
  19. Agnes 2.5 Pro Alpha 67%
  20. MiMo-V2.5-Pro 65%
0.0 44.8 89.5

Pass rate · higher is better

Current configurations with a reported Terminal-Bench v2.1 value; missing values are omitted.

How to interpret the result

What do Terminal-Bench v2.1 results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher pass rate is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read Terminal-Bench v2.1 as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationPass rate
1 GPT-5.6 Sol (xhigh)Leader xhigh 90%
2 Claude Opus 5 Adaptive Reasoning, Max Effort 89%
3 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 88%
4 GPT-5.6 Sol max 88%
5 GPT-5.6 Terra max 88%
6 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 88%
7 GPT-5.6 Sol (high) high 87%
8 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 86%
9 GPT-5.6 Sol (medium) medium 86%
10 Kimi K3 Published configuration 85%
11 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 85%
12 Grok 4.5 high 82%
13 GPT-5.6 Luna max 81%
14 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 81%
15 GPT-5.6 Terra (xhigh) xhigh 80%
16 Gemini 3.5 Flash (high) high 79%
17 GLM-5.2 (max) max 78%
18 GPT-5.6 Luna (xhigh) xhigh 78%
19 Muse Spark 1.1 (xhigh) xhigh 78%
20 Gemini 3.6 Flash (high) high 78%
21 GPT-5.6 Sol (low) low 77%
22 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 76%
23 GPT-5.6 Terra (high) high 76%
24 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 75%
25 Qwen3.7 Max Published configuration 75%
Show the remaining 106 configurations
#ModelConfigurationPass rate
26 GPT-5.6 Sol (Non-reasoning) Non-reasoning 74%
27 Gemini 3.1 Pro (Preview) Published configuration 74%
28 GPT-5.6 Terra (medium) medium 72%
29 Motif 3 (Beta) Beta 71%
30 KAT Coder Pro V2 Published configuration 70%
31 GPT-5.6 Luna (high) high 70%
32 Nex-N2-Pro Published configuration 68%
33 Kimi K2.7 Code Published configuration 67%
34 Agnes 2.5 Pro Alpha Published configuration 67%
35 MiMo-V2.5-Pro Published configuration 65%
36 MiniMax-M3 Published configuration 65%
37 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 65%
38 Hy3 Published configuration 64%
39 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 64%
40 MiMo-V2.5 Published configuration 64%
41 GPT-5.6 Terra (low) low 63%
42 Muse Spark Published configuration 62%
43 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 62%
44 MiMo-V2-Flash (Non-reasoning) Non-reasoning 62%
45 Qwen3.6 Plus Published configuration 61%
46 Qwen3.7 Plus Published configuration 61%
47 Qwen3.6 27B (Reasoning) Reasoning 61%
48 JT-4.1 Flash 236B A21B Published configuration 60%
49 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 57%
50 GPT-5.6 Terra (Non-reasoning) Non-reasoning 56%
51 Inkling (xhigh) xhigh 55%
52 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 54%
53 Gemini 3.5 Flash-Lite Published configuration 54%
54 GPT-5.6 Luna (medium) medium 53%
55 GLM-5.2 (Non-reasoning) Non-reasoning 52%
56 Qwen3.5 397B A17B (Reasoning) Reasoning 51%
57 Qwen3.6 27B (Non-reasoning) Non-reasoning 51%
58 Mistral Medium 3.5 Published configuration 51%
59 LongCat 2.0 Published configuration 50%
60 Qwen3.5 122B A10B (Reasoning) Reasoning 48%
61 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 47%
62 Qwen3.6 35B A3B (Reasoning) Reasoning 45%
63 Claude 4.5 Haiku (Reasoning) Reasoning 44%
64 GPT-5.6 Luna (low) low 43%
65 Gemma 4 31B (Reasoning) Reasoning 43%
66 Ring-2.6-1T Published configuration 43%
67 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 42%
68 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 41%
69 Step 3.7 Flash Published configuration 39%
70 GPT-5.6 Luna (Non-reasoning) Non-reasoning 39%
71 Gemma 4 26B A4B (Reasoning) Reasoning 39%
72 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 39%
73 Qwen3 Coder Next Published configuration 38%
74 North Mini Code Published configuration 36%
75 GPT-5.5 Instant (June 2026) June 2026 35%
76 Grok 4.3 (Non-reasoning) Non-reasoning 34%
77 Gemini 3.1 Flash-Lite Published configuration 31%
78 Devstral 2 Published configuration 30%
79 K-EXAONE (Reasoning) Reasoning 30%
80 Devstral Small 2 Published configuration 30%
81 Nova 2.0 Pro Preview (medium) medium 30%
82 Gemma 4 31B (Non-reasoning) Non-reasoning 29%
83 Qwen3.5 9B (Reasoning) Reasoning 29%
84 Gemini 2.5 Pro Published configuration 28%
85 Gemma 4 12B (Reasoning) Reasoning 27%
86 Mercury 2 Published configuration 27%
87 gpt-oss-120b (high) high 26%
88 Qwen3.5 4B (Reasoning) Reasoning 26%
89 Ling 2.6 Flash Published configuration 24%
90 Command A+ Published configuration 23%
91 EXAONE 4.5 33B Published configuration 21%
92 Qwen3.5 4B (Non-reasoning) Non-reasoning 21%
93 Qwen3.5 9B (Non-reasoning) Non-reasoning 21%
94 Mistral Small 4 (Reasoning) Reasoning 21%
95 Nemotron Cascade 2 30B A3B Published configuration 21%
96 Trinity Large Thinking Published configuration 21%
97 Nova 2.0 Pro Preview (low) low 19%
98 HyperNova 60B 2605 Published configuration 18%
99 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 17%
100 Nova 2.0 Lite (high) high 16%
101 K2 Think V2 Published configuration 15%
102 gpt-oss-120b (low) low 14%
103 gpt-oss-20b (high) high 14%
104 DiffusionGemma 26B A4B Published configuration 12%
105 Magistral Medium 1.2 Published configuration 12%
106 Mistral Large 3 Published configuration 12%
107 Solar Pro 3 Published configuration 12%
108 Ministral 3 14B Published configuration 10%
109 Llama 4 Maverick Published configuration 8%
110 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 7%
111 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 7%
112 Qwen3 Next 80B A3B (Reasoning) Reasoning 7%
113 G9v3-3B Published configuration 6%
114 Llama 3.3 Instruct 70B Published configuration 5%
115 Magistral Small 1.2 Published configuration 4%
116 Ministral 3 8B Published configuration 4%
117 Llama 4 Scout Published configuration 4%
118 NVIDIA Nemotron 3 Nano 4B Published configuration 4%
119 Granite 4.1 8B Published configuration 3%
120 Qwen3.5 2B (Reasoning) Reasoning 3%
121 Granite 4.1 30B Published configuration 3%
122 Gemma 4 E4B (Reasoning) Reasoning 2%
123 Granite 4.1 3B Published configuration 1%
124 Nanbeige4.1-3B Published configuration 1%
125 Gemma 4 E2B (Reasoning) Reasoning 0%
126 Phi-4 Mini Instruct Published configuration 0%
127 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 0%
128 MiniCPM-V 4.6 1.3B Published configuration 0%
129 Ministral 3 3B Published configuration 0%
130 Qwen3.5 0.8B (Reasoning) Reasoning 0%
131 Qwen3.5 2B (Non-reasoning) Non-reasoning 0%

Showing the top 25 of 131 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A terminal-based agent benchmark covering software engineering, system administration, data processing, model training, and security tasks.

How to read the score

The published metric is Pass rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same terminalbenchV21 field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official Terminal-Bench v2.1 resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗