Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

Terminal-Bench Hard benchmark: model scores and methodology

A legacy terminal-use evaluation of agentic coding, system administration, data processing, and related command-line tasks.

Current published leader
GPT-5.6 Sol · max
Top score
66%
Primary metric
Pass rate · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

Published Terminal-Bench Hard model scores in this snapshot

214 configurations · 160 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 214 current configurations with a reported Terminal-Bench Hard score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on Terminal-Bench Hard

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. GPT-5.6 Sol 66%
  2. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 63%
  3. GPT-5.6 Terra xhigh 63%
  4. Gemini 3.1 Pro (Preview) 54%
  5. GPT-5.3 Codex xhigh 53%
  6. GLM-5.2 max 51%
  7. Qwen3.7 Max 51%
  8. KAT Coder Pro V2 49%
  9. Qwen3.7 Plus 47%
  10. DeepSeek V4 Pro Reasoning, Max Effort 46%
  11. Gemini 3.5 Flash minimal 46%
  12. Muse Spark 45%
  13. Kimi K2.7 Code 45%
  14. Qwen3.6 Plus 44%
  15. MiMo-V2.5-Pro 43%
  16. Claude Sonnet 4.6 Non-reasoning, Low Effort 42%
  17. MiniMax-M3 42%
  18. MiMo-V2.5 42%
  19. Qwen3.5 397B A17B Reasoning 41%
  20. DeepSeek V4 Flash Reasoning, High Effort 39%
0.0 33.0 65.9

Pass rate · higher is better

Current configurations with a reported Terminal-Bench Hard value; missing values are omitted.

How to interpret the result

What do Terminal-Bench Hard results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher pass rate is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read Terminal-Bench Hard as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationPass rate
1 GPT-5.6 SolLeader max 66%
2 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 63%
3 GPT-5.6 Sol (medium) medium 63%
4 GPT-5.6 Terra (xhigh) xhigh 63%
5 GPT-5.6 Sol (high) high 62%
6 GPT-5.6 Sol (xhigh) xhigh 61%
7 GPT-5.6 Sol (low) low 61%
8 GPT-5.6 Terra max 58%
9 GPT-5.6 Terra (high) high 58%
10 Gemini 3.1 Pro (Preview) Published configuration 54%
11 GPT-5.3 Codex (xhigh) xhigh 53%
12 GLM-5.2 (max) max 51%
13 Qwen3.7 Max Published configuration 51%
14 KAT Coder Pro V2 Published configuration 49%
15 Qwen3.7 Plus Published configuration 47%
16 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 46%
17 Gemini 3.5 Flash (minimal) minimal 46%
18 Muse Spark Published configuration 45%
19 Kimi K2.7 Code Published configuration 45%
20 GPT-5.6 Terra (low) low 44%
21 Qwen3.6 Plus Published configuration 44%
22 MiMo-V2.5-Pro Published configuration 43%
23 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 42%
24 MiniMax-M3 Published configuration 42%
25 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 42%
Show the remaining 189 configurations
#ModelConfigurationPass rate
26 MiMo-V2.5 Published configuration 42%
27 Gemini 3.5 Flash (high) high 41%
28 Qwen3.5 397B A17B (Reasoning) Reasoning 41%
29 Gemini 3.5 Flash (medium) medium 39%
30 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 39%
31 Kimi K2.6 (Non-reasoning) Non-reasoning 38%
32 o3 Published configuration 37%
33 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 36%
34 Gemma 4 31B (Reasoning) Reasoning 36%
35 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 36%
36 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 36%
37 MiMo-V2-Omni-0327 Published configuration 36%
38 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 36%
39 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 36%
40 Step 3.7 Flash Published configuration 36%
41 MiMo-V2-Omni Published configuration 35%
42 Qwen3.6 27B (Reasoning) Reasoning 35%
43 Qwen3.6 35B A3B (Reasoning) Reasoning 35%
44 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 34%
45 Hy3-preview (Reasoning) Reasoning 34%
46 Mistral Medium 3.5 Published configuration 33%
47 Hy3-preview (Non-reasoning) Non-reasoning 32%
48 Ling-2.6-1T Published configuration 31%
49 MiMo-V2-Flash (Feb 2026) Feb 2026 31%
50 North Mini Code Published configuration 31%
51 Qwen3.5 122B A10B (Reasoning) Reasoning 31%
52 Gemma 4 31B (Non-reasoning) Non-reasoning 30%
53 Grok 4.3 (medium) medium 30%
54 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 30%
55 JT-35B-Flash Published configuration 29%
56 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 29%
57 Ring-2.6-1T Published configuration 29%
58 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 27%
59 Claude 4.5 Haiku (Reasoning) Reasoning 27%
60 Doubao Seed Code Published configuration 27%
61 Gemini 2.5 Pro Published configuration 27%
62 Grok 4.3 (low) low 27%
63 Mercury 2 Published configuration 27%
64 MiMo-V2-Flash (Non-reasoning) Non-reasoning 26%
65 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 26%
66 Command A+ Published configuration 25%
67 ERNIE 5.0 Thinking Preview Published configuration 25%
68 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 25%
69 Gemini 3.1 Flash-Lite Published configuration 24%
70 Nova 2.0 Pro Preview (medium) medium 24%
71 Qwen3.5 9B (Reasoning) Reasoning 24%
72 HyperNova 60B 2605 Published configuration 23%
73 gpt-oss-120b (high) high 23%
74 K-EXAONE (Reasoning) Reasoning 23%
75 Trinity Large Thinking Published configuration 23%
76 Ling 2.6 Flash Published configuration 21%
77 Nemotron Cascade 2 30B A3B Published configuration 21%
78 Qwen3.5 Omni Plus Published configuration 21%
79 Qwen3.6 27B (Non-reasoning) Non-reasoning 21%
80 EXAONE 4.5 33B Published configuration 20%
81 Devstral 2 Published configuration 19%
82 Grok 4.3 (Non-reasoning) Non-reasoning 19%
83 Gemma 4 12B (Reasoning) Reasoning 18%
84 JT-MINI Published configuration 18%
85 Qwen3 Coder Next Published configuration 18%
86 Qwen3.5 4B (Reasoning) Reasoning 18%
87 Qwen3.5 9B (Non-reasoning) Non-reasoning 18%
88 Mistral Small 4 (Reasoning) Reasoning 17%
89 Nova 2.0 Lite (medium) medium 17%
90 Nova 2.0 Pro Preview (low) low 17%
91 Cogito v2.1 (Reasoning) Reasoning 17%
92 Devstral Small 2 Published configuration 17%
93 Nova 2.0 Lite (high) high 17%
94 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 17%
95 Mistral Large 3 Published configuration 16%
96 Apriel-v1.6-15B-Thinker Published configuration 14%
97 Gemma 4 26B A4B (Reasoning) Reasoning 14%
98 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 14%
99 Magistral Medium 1.2 Published configuration 13%
100 HyperCLOVA X SEED Think (32B) 32B 12%
101 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 12%
102 Gemma 4 12B (Non-reasoning) Non-reasoning 11%
103 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 11%
104 Kimi Linear 48B A3B Instruct Published configuration 11%
105 Qwen3.5 4B (Non-reasoning) Non-reasoning 11%
106 LongCat Flash Lite Published configuration 11%
107 Mistral Small 4 (Non-reasoning) Non-reasoning 11%
108 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 11%
109 gpt-oss-20b (high) high 11%
110 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 10%
111 K2-V2 (high) high 10%
112 Qwen3 Next 80B A3B (Reasoning) Reasoning 10%
113 INTELLECT-3 Published configuration 9%
114 KAT-Coder-Pro V1 Published configuration 9%
115 Gemma 4 E4B (Reasoning) Reasoning 8%
116 K2-V2 (medium) medium 8%
117 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 8%
118 Qwen3.5 Omni Flash Published configuration 8%
119 Gemma 4 E4B (Non-reasoning) Non-reasoning 8%
120 Qwen3 Next 80B A3B Instruct Published configuration 8%
121 Ring-flash-2.0 Published configuration 8%
122 Solar Pro 3 Published configuration 8%
123 K-EXAONE (Non-reasoning) Non-reasoning 7%
124 K2 Think V2 Published configuration 7%
125 Llama 3.1 Instruct 405B Published configuration 7%
126 Llama 4 Maverick Published configuration 7%
127 NVIDIA Nemotron 3 Nano 4B Published configuration 7%
128 Nova 2.0 Lite (Non-reasoning) Non-reasoning 7%
129 Nova 2.0 Omni (Non-reasoning) Non-reasoning 7%
130 Nova Premier Published configuration 7%
131 ERNIE 4.5 300B A47B Published configuration 6%
132 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 5%
133 Step3 VL 10B Published configuration 5%
134 gpt-oss-120b (low) low 5%
135 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 5%
136 K2-V2 (low) low 5%
137 LFM2.5-8B-A1B Published configuration 5%
138 Llama 3.1 Nemotron Instruct 70B Published configuration 5%
139 Magistral Small 1.2 Published configuration 5%
140 Ministral 3 14B Published configuration 5%
141 Ministral 3 8B Published configuration 5%
142 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 5%
143 Nova 2.0 Omni (medium) medium 5%
144 gpt-oss-20b (low) low 5%
145 EXAONE 4.0 32B (Reasoning) Reasoning 4%
146 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 4%
147 Motif-2-12.7B-Reasoning Published configuration 4%
148 Nova 2.0 Lite (low) low 4%
149 Nova 2.0 Omni (low) low 4%
150 Phi-4 Published configuration 4%
151 Qwen3 Omni 30B A3B (Reasoning) Reasoning 4%
152 Qwen3.5 2B (Non-reasoning) Non-reasoning 4%
153 Qwen3.5 2B (Reasoning) Reasoning 4%
154 Gemma 4 E2B (Reasoning) Reasoning 3%
155 Llama 3.3 Instruct 70B Published configuration 3%
156 Mi:dm K 2.5 Pro Preview Published configuration 3%
157 Falcon-H1R-7B Published configuration 2%
158 Gemma 4 E2B (Non-reasoning) Non-reasoning 2%
159 Granite 4.0 H Small Published configuration 2%
160 Granite 4.1 30B Published configuration 2%
161 Granite 4.1 3B Published configuration 2%
162 Jamba 1.7 Large Published configuration 2%
163 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 2%
164 Mi:dm K 2.5 Pro Published configuration 2%
165 Sarvam 30B (high) high 2%
166 Solar Open 100B (Reasoning) Reasoning 2%
167 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 2%
168 Granite 4.0 Micro Published configuration 2%
169 Llama 4 Scout Published configuration 2%
170 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 2%
171 Nova Micro Published configuration 2%
172 Qwen3 Omni 30B A3B Instruct Published configuration 2%
173 Sarvam 105B (high) high 2%
174 Command A Published configuration 1%
175 Jamba Reasoning 3B Published configuration 1%
176 LFM2 2.6B Published configuration 1%
177 Ling-mini-2.0 Published configuration 1%
178 Llama 3.2 Instruct 11B (Vision) Vision 1%
179 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 1%
180 Olmo 3 7B Think Published configuration 1%
181 Tri-21B-Think Published configuration 1%
182 Apertus 70B Instruct Published configuration 0%
183 Apertus 8B Instruct Published configuration 0%
184 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 0%
185 Exaone 4.0 1.2B (Reasoning) Reasoning 0%
186 Gemma 3 270M Published configuration 0%
187 Granite 4.0 1B Published configuration 0%
188 Granite 4.0 350M Published configuration 0%
189 Granite 4.0 H 1B Published configuration 0%
190 Granite 4.0 H 350M Published configuration 0%
191 Granite 4.1 8B Published configuration 0%
192 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 0%
193 Jamba 1.7 Mini Published configuration 0%
194 LFM2 24B A2B Published configuration 0%
195 LFM2 8B A1B Published configuration 0%
196 LFM2.5-1.2B-Instruct Published configuration 0%
197 LFM2.5-1.2B-Thinking Published configuration 0%
198 LFM2.5-VL-1.6B Published configuration 0%
199 MiniCPM-V 4.6 1.3B Published configuration 0%
200 MiniCPM5-1B (Non-reasoning) Non-reasoning 0%
201 MiniCPM5-1B (Reasoning) Reasoning 0%
202 Ministral 3 3B Published configuration 0%
203 Molmo 7B-D Published configuration 0%
204 Molmo2-8B Published configuration 0%
205 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 0%
206 Nanbeige4.1-3B Published configuration 0%
207 Olmo 3 7B Instruct Published configuration 0%
208 Olmo 3.1 32B Instruct Published configuration 0%
209 Olmo 3.1 32B Think Published configuration 0%
210 Phi-4 Mini Instruct Published configuration 0%
211 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 0%
212 Qwen3.5 0.8B (Reasoning) Reasoning 0%
213 Reka Flash 3 Published configuration 0%
214 Tiny Aya Global Published configuration 0%

Showing the top 25 of 214 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A legacy terminal-use evaluation of agentic coding, system administration, data processing, and related command-line tasks. Artificial Analysis marks this track as superseded by Terminal-Bench v2.1; historical scores remain useful only within the legacy setup.

How to read the score

The published metric is Pass rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same terminalbenchHard field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official Terminal-Bench Hard resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗