Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of the hardest Terminal-Bench tasks completed, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
282 configurations of 190 models
Unit
% of tasks
Licence
Artificial Analysis commercial data licence

Terminal-Bench Hard: Qwen3.7-Max

Top 15 of 190 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 GPT-5.6 Sol (max reasoning)OpenAI 65.9%
  2. 2 Claude Fable 5 (max reasoning)Anthropic 62.9%
  3. 2 GPT-5.6 Terra (extra-high reasoning)OpenAI 62.9%
  4. 4 GPT-5.5 (extra-high reasoning)OpenAI 60.6%
  5. 5 Claude Opus 4.8 (max reasoning)Anthropic 58.3%
  6. 6 GPT-5.4 (extra-high reasoning)OpenAI 57.6%
  7. 7 Claude Opus 4.7 (no reasoning)Anthropic 54.5%
  8. 8 Gemini 3.1 Pro PreviewGoogle 53.8%
  9. 9 Claude Sonnet 4.6 (max reasoning)Anthropic 53.0%
  10. 9 GPT-5.3-Codex (extra-high reasoning)OpenAI 53.0%
  11. 11 GPT-5.4 mini (extra-high reasoning)OpenAI 52.3%
  12. 12 Qwen3.7-MaxAlibaba 50.8%
  13. 12 GLM-5.2 (max reasoning)Z.ai 50.8%
  14. 14 Claude Opus 4.6 (no reasoning)Anthropic 48.5%
  15. 15 Qwen3.7-PlusAlibaba 47.0%

Full results

Artificial Analysis: Terminal-Bench Hard, % of tasks, higher is better
#ModelTerminal-Bench Hard
% of tasks, higher is better
Price
$ per million tokens, in / out
1 GPT-5.6 Sol (max reasoning)OpenAI · best of 5 settings
65.9%
$4 / $20
2 Claude Fable 5 (max reasoning)Anthropic
62.9%
$10 / $50
2 GPT-5.6 Terra (extra-high reasoning)OpenAI · best of 4 settings
62.9%
$2 / $12
4 GPT-5.5 (extra-high reasoning)OpenAI · best of 5 settings
60.6%
$5 / $30
5 Claude Opus 4.8 (max reasoning)Anthropic
58.3%
$5 / $25
6 GPT-5.4 (extra-high reasoning)OpenAI · best of 3 settings
57.6%
$2.50 / $15
7 Claude Opus 4.7 (no reasoning)Anthropic · best of 2 settings
54.5%
$5 / $25
8 Gemini 3.1 Pro PreviewGoogle
53.8%
$2 / $12
9 Claude Sonnet 4.6 (max reasoning)Anthropic · best of 2 settings
53.0%
$3 / $15
9 GPT-5.3-Codex (extra-high reasoning)OpenAI
53.0%
$1.75 / $14
11 GPT-5.4 mini (extra-high reasoning)OpenAI · best of 3 settings
52.3%
$0.75 / $4.50
12 Qwen3.7-MaxAlibaba
50.8%
$1.48 / $4.42
12 GLM-5.2 (max reasoning)Z.ai
50.8%
$1.40 / $4.40
14 Claude Opus 4.6 (no reasoning)Anthropic · best of 2 settings
48.5%
$5 / $25
15 Qwen3.7-PlusAlibaba
47.0%
$0.32 / $1.28
15 Claude Opus 4.5 (reasoning on)Anthropic · best of 2 settings
47.0%
$5 / $25
15 GPT-5.2 (extra-high reasoning)OpenAI · best of 3 settings
47.0%
$1.75 / $14
18 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek · best of 3 settings
46.2%
$1.42 / $2.83
18 Gemini 3.5 Flash (minimal reasoning)Google · best of 3 settings
46.2%
$1.50 / $9
20 Muse SparkMeta
45.5%
–
20 GPT-5.1 (high reasoning)OpenAI · best of 2 settings
45.5%
$1.25 / $10
22 Kimi K2.7 CodeMoonshot AI
44.7%
$0.95 / $4
23 Qwen3.6-Max-PreviewAlibaba
43.9%
$1.03 / $6.16
23 Qwen3.6-PlusAlibaba
43.9%
$0.33 / $1.95
23 Kimi K2.6Moonshot AI · best of 2 settings
43.9%
$0.95 / $4
26 MiMo-V2.5-Pro (reasoning on)Xiaomi · best of 2 settings
43.2%
$0.43 / $0.87
26 GLM-5.1Z.ai · best of 2 settings
43.2%
$1.38 / $4.40
26 GLM-5Z.ai · best of 2 settings
43.2%
$0.95 / $2.55
29 MiniMax-M3MiniMax
42.4%
$0.30 / $1.20
29 GPT-5.4 nano (extra-high reasoning)OpenAI · best of 3 settings
42.4%
$0.20 / $1.25
29 GPT-5.5 Instant (2026-05-26)OpenAI
42.4%
–
32 Gemini 3 Pro Preview (high reasoning)Google · best of 2 settings
41.7%
–
32 MiMo-V2.5Xiaomi
41.7%
$0.17 / $0.34
34 Qwen3.5-397B-A17BAlibaba · best of 2 settings
40.9%
$0.55 / $3.50
34 Grok 4.20 (0309, reasoning)xAI
40.9%
–
36 MiniMax-M2.7MiniMax
39.4%
$0.30 / $1.20
37 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek · best of 3 settings
38.6%
$0.14 / $0.28
37 Gemini 3 Flash Preview (reasoning on)Google · best of 2 settings
38.6%
$0.50 / $3
39 GPT-5-Codex (high reasoning)OpenAI
37.9%
–
39 GPT-5 (medium reasoning)OpenAI · best of 4 settings
37.9%
$1.25 / $10
39 Grok 4.20xAI · best of 2 settings
37.9%
$1.25 / $2.50
39 Grok 4.3 (high reasoning)xAI · best of 4 settings
37.9%
$1.25 / $2.50
39 Grok 4xAI
37.9%
–
44 GPT-5.2-Codex (extra-high reasoning)OpenAI
37.1%
$1.75 / $14
44 o3OpenAI
37.1%
$2 / $8
46 Gemma 4 31BGoogle · best of 2 settings
36.4%
$0.14 / $0.40
47 Claude Sonnet 4.5 (reasoning on)Anthropic · best of 2 settings
35.6%
$3 / $15
47 DeepSeek-V3.2 (reasoning on)DeepSeek · best of 2 settings
35.6%
$0.30 / $0.96
49 Qwen3.6-27BAlibaba · best of 2 settings
34.8%
$0.30 / $3.20
49 Qwen3.6-35B-A3BAlibaba · best of 2 settings
34.8%
$0.10 / $1
49 DeepSeek-V3.2-SpecialeDeepSeek
34.8%
–
49 MiniMax-M2.5MiniMax
34.8%
$0.30 / $1.20
49 Kimi K2.5Moonshot AI · best of 2 settings
34.8%
$0.57 / $2.85
49 GPT-5.1-Codex (high reasoning)OpenAI
34.8%
$1.25 / $10
55 Claude Opus 4.1 (reasoning on)Anthropic
34.3%
$15 / $75
56 Hy3 Preview (reasoning on)Tencent · best of 2 settings
34.1%
$0.18 / $0.60
57 Mistral Medium 3.5Mistral AI
33.3%
$1.50 / $7.50
57 GPT-5.1-Codex-Mini (high reasoning)OpenAI
33.3%
$0.25 / $2
57 GPT-5 mini (high reasoning)OpenAI · best of 3 settings
33.3%
$0.25 / $2
57 GLM-5-TurboZ.ai
33.3%
$1.20 / $4
61 Qwen3.5-27BAlibaba · best of 2 settings
32.6%
$0.27 / $2.16
61 GLM-5V-TurboZ.ai
32.6%
$1.20 / $4
63 DeepSeek-V3.1-Terminus (no reasoning)DeepSeek · best of 2 settings
31.8%
$0.27 / $1
63 GLM-4.7Z.ai · best of 2 settings
31.8%
$0.54 / $1.98
65 Qwen3.5-122B-A10BAlibaba · best of 2 settings
31.1%
$0.26 / $2.08
65 Claude Sonnet 4 (reasoning on)Anthropic · best of 2 settings
31.1%
$3 / $15
65 DeepSeek-V3.2-Exp (reasoning on)DeepSeek · best of 2 settings
31.1%
$0.27 / $0.41
65 Kimi K2 ThinkingMoonshot AI
31.1%
$0.60 / $2.50
69 MiniMax-M2.1MiniMax
28.8%
$0.30 / $1.20
69 Nemotron 3 Super 120B A12BNVIDIA
28.8%
$0.085 / $0.40
69 GLM-4.6 (no reasoning)Z.ai · best of 2 settings
28.8%
$0.50 / $2
72 Claude Haiku 4.5 (no reasoning)Anthropic · best of 2 settings
27.3%
$1 / $5
72 Step 3.5 FlashStepFun
27.3%
$0.10 / $0.30
74 Qwen3.5-35B-A3BAlibaba · best of 2 settings
26.5%
$0.16 / $1.30
74 Gemini 2.5 ProGoogle
26.5%
$1.25 / $10
74 Mercury 2Inception
26.5%
$0.25 / $0.75
77 MiniMax-M2MiniMax
25.8%
$0.30 / $1.20
78 DeepSeek-V3.1 (reasoning on)DeepSeek · best of 2 settings
25.0%
$0.55 / $1.65
78 Gemma 4 26B A4B (no reasoning)Google · best of 2 settings
25.0%
$0.10 / $0.30
80 Qwen3.5-9BAlibaba · best of 2 settings
24.2%
$0.10 / $0.15
80 Qwen3-Max-ThinkingAlibaba
24.2%
$0.78 / $3.90
80 Gemini 3.1 Flash-Lite PreviewGoogle
24.2%
$0.25 / $1.50
80 Grok 4.1 Fast (reasoning)xAI
24.2%
–
84 Kimi K2 (0905)Moonshot AI
23.5%
$0.60 / $2.50
84 gpt-oss-120b (high reasoning)OpenAI · best of 2 settings
23.5%
$0.15 / $0.60
86 Trinity Large ThinkingArcee AI
22.7%
$0.25 / $0.80
87 Grok 4.20 (0309, non-reasoning, no reasoning)xAI
22.0%
–
87 GLM-4.5Z.ai
22.0%
$0.60 / $2.20
87 GLM-4.7-FlashZ.ai · best of 2 settings
22.0%
$0.06 / $0.40
90 Qwen3.5-Omni-PlusAlibaba
21.2%
–
91 Qwen3-MaxAlibaba
20.5%
$0.78 / $3.90
91 GLM-4.5-AirZ.ai
20.5%
$0.14 / $0.86
93 Qwen3-Max-PreviewAlibaba
19.7%
–
94 Qwen3-Coder-480B-A35BAlibaba
18.9%
$0.35 / $1.50
94 Devstral 2Mistral AI
18.9%
$0.40 / $2
94 Grok 4 Fast (reasoning)xAI
18.9%
–
97 Qwen3.5-4B (reasoning on)Alibaba · best of 2 settings
18.2%
–
97 Qwen3-Coder-NextAlibaba
18.2%
$0.18 / $0.90
97 Gemma 4 12BGoogle · best of 2 settings
18.2%
–
100 Qwen3-Max-Thinking-PreviewAlibaba
17.4%
–
100 Nova 2 Lite (medium reasoning)Amazon · best of 4 settings
17.4%
$0.30 / $2.50
100 Mistral Small 4Mistral AI · best of 2 settings
17.4%
$0.15 / $0.60
100 GPT-5 nano (medium reasoning)OpenAI · best of 3 settings
17.4%
$0.05 / $0.40
100 Grok Code Fast 1xAI
17.4%
–
105 Gemini 2.5 Flash Preview (09-2025, reasoning on)Google · best of 2 settings
16.7%
–
105 Devstral Small 2Mistral AI
16.7%
–
107 DeepSeek-R1-0528DeepSeek
15.9%
$0.50 / $2.18
107 Mistral Large 3Mistral AI
15.9%
$0.50 / $1.50
107 Kimi K2 (0711)Moonshot AI
15.9%
$0.57 / $2.30
110 Qwen3-235B-A22B-Instruct-2507Alibaba
15.2%
$0.15 / $0.75
110 Qwen3-Coder-30B-A3B-InstructAlibaba
15.2%
$0.07 / $0.28
110 DeepSeek-V3-0324DeepSeek
15.2%
$0.25 / $1
110 o4-mini (high reasoning)OpenAI
15.2%
$1.10 / $4.40
114 Grok 4.1 Fast (non-reasoning, no reasoning)xAI
14.4%
–
114 GLM-4.6V (reasoning on)Z.ai · best of 2 settings
14.4%
$0.30 / $0.90
116 Qwen3-235B-A22B-Thinking-2507Alibaba
13.6%
$0.30 / $3
116 Gemini 2.5 Flash (reasoning on)Google · best of 2 settings
13.6%
$0.30 / $2.50
116 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA · best of 2 settings
13.6%
$0.05 / $0.20
116 GPT-4.1OpenAI
13.6%
$2 / $8
120 Gemini 2.5 Flash-Lite Preview (09-2025, reasoning on)Google · best of 2 settings
12.9%
–
120 Magistral Medium 1.2Mistral AI
12.9%
–
120 GPT-5 ChatOpenAI
12.9%
–
120 o1OpenAI
12.9%
$15 / $60
124 Grok 4 Fast (non-reasoning, no reasoning)xAI
12.1%
–
125 Qwen3-VL-235B-A22B-ThinkingAlibaba
11.4%
$0.40 / $4
125 Kimi Linear 48B A3B InstructMoonshot AI
11.4%
–
127 Mistral Medium 3.1Mistral AI
10.6%
$0.40 / $2
127 gpt-oss-20b (high reasoning)OpenAI · best of 2 settings
10.6%
$0.03 / $0.15
129 Qwen3-Next-80B-A3B-ThinkingAlibaba
9.8%
$0.15 / $1.20
130 Devstral MediumMistral AI
9.1%
–
131 Qwen3.5-Omni-FlashAlibaba
8.3%
–
131 Qwen3-VL-32B-InstructAlibaba
8.3%
$0.10 / $0.42
131 Gemma 4 E4BGoogle · best of 2 settings
8.3%
–
131 GPT-4o (2024-08-06)OpenAI
8.3%
$2.50 / $10
135 Qwen3-Next-80B-A3B-InstructAlibaba
7.6%
$0.10 / $1.10
135 Qwen3-VL-32B-ThinkingAlibaba
7.6%
–
135 Mistral Small 3.1 24BMistral AI
7.6%
$0.35 / $0.56
135 GPT-4.1 miniOpenAI
7.6%
$0.40 / $1.60
135 Solar Pro 3Upstage
7.6%
$0.15 / $0.60
140 Qwen3-30B-A3B (no reasoning)Alibaba · best of 2 settings
6.8%
$0.12 / $0.50
140 Qwen3-VL-235B-A22B-InstructAlibaba
6.8%
$0.30 / $1.50
140 DeepSeek-V3DeepSeek
6.8%
$0.26 / $1.03
140 Llama 4 MaverickMeta
6.8%
$0.27 / $0.85
140 Mistral Small 3.2 24BMistral AI
6.8%
$0.094 / $0.25
140 o3-miniOpenAI · best of 2 settings
6.8%
$1.10 / $4.40
140 GLM-4.5V (no reasoning)Z.ai · best of 2 settings
6.8%
$0.60 / $1.80
147 Qwen3-235B-A22B (no reasoning)Alibaba · best of 2 settings
6.1%
$0.46 / $1.82
147 Qwen3-30B-A3B-Instruct-2507Alibaba
6.1%
$0.09 / $0.30
147 Qwen3-VL-30B-A3B-InstructAlibaba
6.1%
$0.15 / $0.60
147 Nova Pro 1.0Amazon
6.1%
$0.80 / $3.20
147 DeepSeek-R1DeepSeek
6.1%
$0.70 / $2.50
147 Devstral Small 1.1Mistral AI
6.1%
–
153 Qwen3-14B (no reasoning)Alibaba · best of 2 settings
5.3%
$0.12 / $0.24
153 Qwen3-30B-A3B-Thinking-2507Alibaba
5.3%
$0.20 / $2.40
153 Qwen3-VL-30B-A3B-ThinkingAlibaba
5.3%
$0.29 / $1
156 Qwen2.5-72B-InstructAlibaba
4.5%
$0.36 / $0.40
156 Qwen3-4B-Instruct-2507Alibaba
4.5%
–
156 Gemini 2.5 Flash-Lite (reasoning on)Google · best of 2 settings
4.5%
$0.10 / $0.40
156 Magistral Small 1.2Mistral AI
4.5%
–
156 Ministral 3 14BMistral AI
4.5%
$0.20 / $0.20
156 Ministral 3 8BMistral AI
4.5%
$0.15 / $0.15
162 Qwen3.5-2B (no reasoning)Alibaba · best of 2 settings
3.8%
–
162 Qwen3-Omni-30B-A3B-ThinkingAlibaba
3.8%
–
162 Qwen3-VL-8B-ThinkingAlibaba
3.8%
$0.18 / $2.10
162 Gemma 3 27BGoogle
3.8%
$0.12 / $0.20
162 Phi-4Microsoft
3.8%
$0.07 / $0.14
162 Mistral Medium 3Mistral AI
3.8%
$0.40 / $2
162 GPT-4.1 nanoOpenAI
3.8%
$0.10 / $0.40
169 Qwen3-32B (reasoning on)Alibaba
3.0%
$0.14 / $0.40
169 Gemma 4 E2BGoogle · best of 2 settings
3.0%
–
169 Llama 3.1 70B InstructMeta
3.0%
$0.40 / $0.40
169 Llama 3.3 70B InstructMeta
3.0%
$0.59 / $0.79
173 Qwen3-8B (no reasoning)Alibaba · best of 2 settings
2.3%
$0.12 / $0.46
173 Qwen3-VL-8B-InstructAlibaba
2.3%
$0.12 / $0.46
175 Qwen3-4B-Thinking-2507Alibaba
1.5%
–
175 Qwen3-Omni-30B-A3B-InstructAlibaba
1.5%
–
175 Qwen3-VL-4B-ThinkingAlibaba
1.5%
–
175 Nova Micro 1.0Amazon
1.5%
$0.035 / $0.14
175 Llama 4 ScoutMeta
1.5%
$0.18 / $0.59
180 Nova Lite 1.0Amazon
0.8%
$0.06 / $0.24
180 Command ACohere
0.8%
$2.50 / $10
180 Gemma 3 12BGoogle
0.8%
$0.05 / $0.15
180 Gemma 3 4BGoogle
0.8%
$0.05 / $0.10
180 Llama 3.1 8B InstructMeta
0.8%
$0.05 / $0.08
185 Qwen3.5-0.8B (no reasoning)Alibaba · best of 2 settings
0.0%
–
185 Qwen3-VL-4B-InstructAlibaba
0.0%
–
185 Gemma 3 270MGoogle
0.0%
–
185 Llama 3.2 1B InstructMeta
0.0%
$0.027 / $0.20
185 Ministral 3 3BMistral AI
0.0%
$0.10 / $0.10
185 Reka Flash 3Reka AI
0.0%
$0.10 / $0.20

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of the hardest Terminal-Bench tasks completed, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not other harnesses; Artificial Analysis no longer runs it on new models.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.