Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of terminal tasks (software, system and data jobs) an agent completes, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
187 configurations of 112 models
Unit
% of tasks
Licence
Artificial Analysis commercial data licence

Terminal-Bench 4.0: GPT-6.1 Sol

Top 16 of 187 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Claude Sonnet 5.5 (max reasoning)Anthropic 63.6%
  2. 2 Claude Opus 5.5 (max reasoning)Anthropic 59.6%
  3. 2 Claude Opus 5.5 (extra-high reasoning)Anthropic 59.6%
  4. 2 GPT-6 Astra (extra-high reasoning)OpenAI 59.6%
  5. 5 GPT-6 Astra (max reasoning)OpenAI 59.1%
  6. 6 Claude Sonnet 5.5 (extra-high reasoning)Anthropic 57.1%
  7. 6 Gemini 4 Argon (high reasoning)Google 57.1%
  8. 8 Claude Opus 5.5 (high reasoning)Anthropic 56.6%
  9. 9 GPT-6.1 Sol (max reasoning)OpenAI 56.1%
  10. 10 Claude Fable 5.1 (extra-high reasoning)Anthropic 55.1%
  11. 11 GPT-6.1 Sol (extra-high reasoning)OpenAI 54.0%
  12. 11 GPT-6 Astra (high reasoning)OpenAI 54.0%
  13. 13 Claude Opus 5.5 (medium reasoning)Anthropic 52.5%
  14. 14 Claude Fable 5.1 (high reasoning)Anthropic 52.0%
  15. 14 Claude Fable 5.1 (max reasoning)Anthropic 52.0%
  16. 16 GPT-6.1 Sol (high reasoning)OpenAI 51.5%
  17. 19 GPT-6.1 Sol (medium reasoning)OpenAI 48.0%
  18. 39 GPT-6.1 Sol (low reasoning)OpenAI 30.8%

Full results

Artificial Analysis: Terminal-Bench 4.0, % of tasks, higher is better
#ModelTerminal-Bench 4.0
% of tasks, higher is better
Price
$ per million tokens, in / out
1 Claude Sonnet 5.5 (max reasoning)Anthropic
63.6%
$2 / $10
2 Claude Opus 5.5 (max reasoning)Anthropic
59.6%
$4 / $20
2 Claude Opus 5.5 (extra-high reasoning)Anthropic
59.6%
$4 / $20
2 GPT-6 Astra (extra-high reasoning)OpenAI
59.6%
$10 / $50
5 GPT-6 Astra (max reasoning)OpenAI
59.1%
$10 / $50
6 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
57.1%
$2 / $10
6 Gemini 4 Argon (high reasoning)Google
57.1%
–
8 Claude Opus 5.5 (high reasoning)Anthropic
56.6%
$4 / $20
9 GPT-6.1 Sol (max reasoning)OpenAI
56.1%
$2 / $10
10 Claude Fable 5.1 (extra-high reasoning)Anthropic
55.1%
$10 / $50
11 GPT-6.1 Sol (extra-high reasoning)OpenAI
54.0%
$2 / $10
11 GPT-6 Astra (high reasoning)OpenAI
54.0%
$10 / $50
13 Claude Opus 5.5 (medium reasoning)Anthropic
52.5%
$4 / $20
14 Claude Fable 5.1 (high reasoning)Anthropic
52.0%
$10 / $50
14 Claude Fable 5.1 (max reasoning)Anthropic
52.0%
$10 / $50
16 GPT-6.1 Sol (high reasoning)OpenAI
51.5%
$2 / $10
17 GPT-6 Astra (medium reasoning)OpenAI
49.5%
$10 / $50
18 Claude Opus 5 (max reasoning)Anthropic
49.0%
$5 / $25
19 GPT-6.1 Sol (medium reasoning)OpenAI
48.0%
$2 / $10
20 Claude Opus 5 (extra-high reasoning)Anthropic
46.5%
$5 / $25
21 Claude Opus 5 (high reasoning)Anthropic
46.0%
$5 / $25
22 Claude Fable 5.1 (medium reasoning)Anthropic
44.9%
$10 / $50
23 Claude Sonnet 5.5 (high reasoning)Anthropic
43.9%
$2 / $10
23 GPT-6 Sol (max reasoning)OpenAI
43.9%
$2 / $10
25 Claude Fable 5 (max reasoning)Anthropic
42.4%
$10 / $50
26 GPT-6 Astra (low reasoning)OpenAI
41.9%
$10 / $50
26 GLM-5.3 (max reasoning)Z.ai
41.9%
$1.40 / $4.40
28 Claude Fable 5.1 (low reasoning)Anthropic
40.4%
$10 / $50
29 GPT-5.6 Sol (max reasoning)OpenAI
39.9%
$4 / $20
30 Qwen3.8-Max (0902)Alibaba
38.9%
$2 / $6
31 GPT-5.6 Terra (max reasoning)OpenAI
35.4%
$2 / $12
32 MiMo-V2.6-ProXiaomi
34.8%
$0.43 / $0.87
32 GLM-5.3 (low reasoning)Z.ai
34.8%
$1.40 / $4.40
34 Claude Opus 5 (medium reasoning)Anthropic
34.3%
$5 / $25
35 Muse Spark 1.3 (max reasoning)Meta
33.3%
$1.25 / $4.25
36 Claude Haiku 5.5 (max reasoning)Anthropic
32.8%
$0.10 / $0.50
36 GLM-5.3-FlashZ.ai
32.8%
$0.15 / $0.50
38 Claude Opus 5.5 (low reasoning)Anthropic
31.3%
$4 / $20
39 GPT-6.1 Sol (low reasoning)OpenAI
30.8%
$2 / $10
40 GPT-6 Sol (extra-high reasoning)OpenAI
30.3%
$2 / $10
41 Claude Sonnet 5.5 (medium reasoning)Anthropic
29.8%
$2 / $10
42 Claude Haiku 5.5 (extra-high reasoning)Anthropic
29.3%
$0.10 / $0.50
43 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
26.8%
$0.30 / $1.20
43 Mistral Large 4Mistral AI
26.8%
$1.36 / $4.18
45 Claude Opus 5 (low reasoning)Anthropic
26.3%
$5 / $25
45 GPT-6 Sol (high reasoning)OpenAI
26.3%
$2 / $10
47 Grok 4.7 (extra-high reasoning)xAI
25.8%
$2 / $6
48 Qwen3.8-Flash-NextAlibaba
25.3%
–
49 GPT-5.6 Sol (extra-high reasoning)OpenAI
24.7%
$4 / $20
49 Grok 4.7 (high reasoning)xAI
24.7%
$2 / $6
51 MiMo-V2.6-FlashXiaomi
22.7%
$0.14 / $0.28
52 Claude Haiku 5.5 (high reasoning)Anthropic
21.7%
$0.10 / $0.50
52 Claude Opus 4.8 (max reasoning)Anthropic
21.7%
$5 / $25
54 Grok 4.6 (high reasoning)xAI
21.2%
$2 / $6
55 Claude Sonnet 5.5 (low reasoning)Anthropic
20.7%
$2 / $10
55 GPT-5.6 Sol (high reasoning)OpenAI
20.7%
$4 / $20
57 Gemini 3.8 Flash (high reasoning)Google
19.7%
$1.50 / $7.50
57 Gemini 3.8 Flash (medium reasoning)Google
19.7%
$1.50 / $7.50
59 Qwen3.8-Max (0803)Alibaba
18.7%
–
59 GPT-6 Sol (medium reasoning)OpenAI
18.7%
$2 / $10
61 Grok 4.6 (extra-high reasoning)xAI
17.2%
$2 / $6
62 Muse Spark 1.3 (extra-high reasoning)Meta
16.7%
$1.25 / $4.25
63 Grok 4.7 (low reasoning)xAI
16.2%
$2 / $6
64 Claude Haiku 5.5 (medium reasoning)Anthropic
15.2%
$0.10 / $0.50
65 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek
14.6%
$1.42 / $2.83
65 GPT-5.5 (extra-high reasoning)OpenAI
14.6%
$5 / $30
65 GPT-5.6 Sol (medium reasoning)OpenAI
14.6%
$4 / $20
68 Claude Sonnet 5 (max reasoning)Anthropic
14.1%
$2 / $10
68 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek
14.1%
$1.32 / $3.96
70 Gemini 3.7 Flash (high reasoning)Google
13.6%
$1.50 / $7.50
71 GPT-6 Sol (no reasoning)OpenAI
13.1%
$2 / $10
71 Grok 4.6 (medium reasoning)xAI
13.1%
$2 / $6
73 Claude Haiku 5.5 (low reasoning)Anthropic
12.6%
$0.10 / $0.50
73 Kimi K3 (low reasoning)Moonshot AI
12.6%
$3 / $15
73 Kimi K3 (max reasoning)Moonshot AI
12.6%
$3 / $15
73 GPT-5.5 Instant (2026-06-26)OpenAI
12.6%
–
73 GPT-6 Luna (max reasoning)OpenAI
12.6%
$0.10 / $0.50
78 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
12.1%
$0.14 / $0.28
78 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
12.1%
$0.44 / $1.32
80 GPT-5.6 Luna (max reasoning)OpenAI
11.6%
$0.20 / $1.20
81 Qwen3.8-2.4T-A95BAlibaba
11.1%
$2 / $6
82 Grok 4.5 (high reasoning)xAI
10.6%
$2 / $6
83 Gemini 3.8 Flash (low reasoning)Google
10.1%
$1.50 / $7.50
83 GPT-5.6 Terra (extra-high reasoning)OpenAI
10.1%
$2 / $12
85 GPT-5.5 (high reasoning)OpenAI
9.1%
$5 / $30
85 GPT-6 Sol (low reasoning)OpenAI
9.1%
$2 / $10
87 GPT-6 Luna (extra-high reasoning)OpenAI
8.1%
$0.10 / $0.50
88 Claude Sonnet 5 (extra-high reasoning)Anthropic
7.1%
$2 / $10
88 Gemini 3.6 Flash (high reasoning)Google
7.1%
$1.50 / $7.50
88 Muse Spark 1.2 (extra-high reasoning)Meta
7.1%
$1.25 / $4.25
91 Gemini 3.5 Flash (high reasoning)Google
6.6%
$1.50 / $9
92 Muse Spark 1.1 (extra-high reasoning)Meta
6.1%
$1.25 / $4.25
93 Qwen3.8-27B (extra-high reasoning)Alibaba
5.6%
$0.50 / $3
93 DeepSeek-V4.1-Flash (no reasoning)DeepSeek
5.6%
$0.30 / $1.20
95 Qwen3.8-27B (medium reasoning)Alibaba
5.1%
$0.50 / $3
95 Claude Sonnet 5 (high reasoning)Anthropic
5.1%
$2 / $10
95 GPT-5.5 (medium reasoning)OpenAI
5.1%
$5 / $30
98 GPT-6 Luna (high reasoning)OpenAI
4.5%
$0.10 / $0.50
99 Gemini 3.1 Pro PreviewGoogle
4.0%
$2 / $12
100 GPT-5.6 Luna (extra-high reasoning)OpenAI
3.5%
$0.20 / $1.20
101 Claude Sonnet 4.6 (max reasoning)Anthropic
3.0%
$3 / $15
101 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
3.0%
$0.14 / $0.28
101 Grok 4.6 (low reasoning)xAI
3.0%
$2 / $6
104 Qwen3.8-27B (low reasoning)Alibaba
2.5%
$0.50 / $3
104 Claude Sonnet 5 (low reasoning)Anthropic
2.5%
$2 / $10
104 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek
2.5%
$0.14 / $0.28
104 GPT-5.6 Luna (high reasoning)OpenAI
2.5%
$0.20 / $1.20
104 GPT-6 Luna (medium reasoning)OpenAI
2.5%
$0.10 / $0.50
109 Claude Sonnet 5 (medium reasoning)Anthropic
2.0%
$2 / $10
109 MiniMax-M3MiniMax
2.0%
$0.30 / $1.20
109 GPT-5.4 mini (extra-high reasoning)OpenAI
2.0%
$0.75 / $4.50
109 GLM-5.1Z.ai
2.0%
$1.38 / $4.40
113 Qwen3.7-MaxAlibaba
1.5%
$1.48 / $4.42
113 DeepSeek-V4-Pro (0813, no reasoning)DeepSeek
1.5%
$1.32 / $3.96
113 GPT-5.6 Terra (high reasoning)OpenAI
1.5%
$2 / $12
113 GPT-5.6 Terra (low reasoning)OpenAI
1.5%
$2 / $12
113 GPT-6 Luna (no reasoning)OpenAI
1.5%
$0.10 / $0.50
118 Qwen3.7-PlusAlibaba
1.0%
$0.32 / $1.28
118 Gemini 3.5 Flash-LiteGoogle
1.0%
$0.30 / $2.50
118 Kimi K2.7 CodeMoonshot AI
1.0%
$0.95 / $4
118 GPT-5.6 Luna (no reasoning)OpenAI
1.0%
$0.20 / $1.20
118 GPT-5.6 Sol (low reasoning)OpenAI
1.0%
$4 / $20
118 GPT-5.6 Terra (medium reasoning)OpenAI
1.0%
$2 / $12
118 Inkling SmallThinking Machines
1.0%
$0.45 / $1.20
118 Inkling (extra-high reasoning)Thinking Machines
1.0%
$0.95 / $4.05
118 GLM-5.2 (max reasoning)Z.ai
1.0%
$1.40 / $4.40
127 Qwen3.5-9BAlibaba
0.5%
$0.10 / $0.15
127 Trinity Large ThinkingArcee AI
0.5%
$0.25 / $0.80
127 Gemini 3.1 Flash-Lite PreviewGoogle
0.5%
$0.25 / $1.50
127 Muse Glimmer (high reasoning)Meta
0.5%
–
127 Kimi K2.6Moonshot AI
0.5%
$0.95 / $4
127 GPT-5.4 nano (extra-high reasoning)OpenAI
0.5%
$0.20 / $1.25
127 GPT-5.6 Luna (medium reasoning)OpenAI
0.5%
$0.20 / $1.20
127 GPT-5.6 Terra (no reasoning)OpenAI
0.5%
$2 / $12
127 Hy3Tencent
0.5%
$0.14 / $0.58
127 Solar Pro 4Upstage
0.5%
$0.09 / $0.36
137 Qwen3-14B (reasoning on)Alibaba
0.0%
$0.12 / $0.24
137 Qwen3-235B-A22B-Thinking-2507Alibaba
0.0%
$0.30 / $3
137 Qwen3-30B-A3B-Thinking-2507Alibaba
0.0%
$0.20 / $2.40
137 Qwen3-32B (reasoning on)Alibaba
0.0%
$0.14 / $0.40
137 Qwen3.5-122B-A10BAlibaba
0.0%
$0.26 / $2.08
137 Qwen3.5-397B-A17BAlibaba
0.0%
$0.55 / $3.50
137 Qwen3.6-27BAlibaba
0.0%
$0.30 / $3.20
137 Qwen3.6-35B-A3BAlibaba
0.0%
$0.10 / $1
137 Qwen3.8-27B (no reasoning)Alibaba
0.0%
$0.50 / $3
137 Qwen3-8B (reasoning on)Alibaba
0.0%
$0.12 / $0.46
137 Qwen3-Coder-NextAlibaba
0.0%
$0.18 / $0.90
137 Claude Haiku 4.5 (reasoning on)Anthropic
0.0%
$1 / $5
137 Claude Sonnet 4.5 (reasoning on)Anthropic
0.0%
$3 / $15
137 DeepSeek-V3-0324DeepSeek
0.0%
$0.25 / $1
137 DeepSeek-V3DeepSeek
0.0%
$0.26 / $1.03
137 DeepSeek-R1DeepSeek
0.0%
$0.70 / $2.50
137 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
0.0%
$0.27 / $1
137 Gemini 2.5 ProGoogle
0.0%
$1.25 / $10
137 Gemma 3 12BGoogle
0.0%
$0.05 / $0.15
137 Gemma 3 27BGoogle
0.0%
$0.12 / $0.20
137 Gemma 4 31BGoogle
0.0%
$0.14 / $0.40
137 Gemma 4 E4BGoogle
0.0%
–
137 Granite 4.2 8BIBM
0.0%
$0.06 / $0.25
137 Mercury 2Inception
0.0%
$0.25 / $0.75
137 Llama 4 MaverickMeta
0.0%
$0.27 / $0.85
137 Llama 4 ScoutMeta
0.0%
$0.18 / $0.59
137 MiniMax-M2.7MiniMax
0.0%
$0.30 / $1.20
137 Devstral 2Mistral AI
0.0%
$0.40 / $2
137 Devstral Small 2Mistral AI
0.0%
–
137 Ministral 3 14BMistral AI
0.0%
$0.20 / $0.20
137 Ministral 3 3BMistral AI
0.0%
$0.10 / $0.10
137 Ministral 3 8BMistral AI
0.0%
$0.15 / $0.15
137 Mistral Large 3Mistral AI
0.0%
$0.50 / $1.50
137 Mistral Medium 3.1Mistral AI
0.0%
$0.40 / $2
137 Mistral Medium 3.5Mistral AI
0.0%
$1.50 / $7.50
137 Mistral Small 4Mistral AI
0.0%
$0.15 / $0.60
137 Mistral Small 3.1 24BMistral AI
0.0%
$0.35 / $0.56
137 Mistral Small 3.2 24BMistral AI
0.0%
$0.094 / $0.25
137 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA
0.0%
$0.05 / $0.20
137 Nemotron 3 Super 120B A12BNVIDIA
0.0%
$0.085 / $0.40
137 GPT-5.6 Luna (low reasoning)OpenAI
0.0%
$0.20 / $1.20
137 GPT-5 mini (high reasoning)OpenAI
0.0%
$0.25 / $2
137 GPT-6 Luna (low reasoning)OpenAI
0.0%
$0.10 / $0.50
137 gpt-oss-120b (high reasoning)OpenAI
0.0%
$0.15 / $0.60
137 gpt-oss-20b (high reasoning)OpenAI
0.0%
$0.03 / $0.15
137 o3-mini (high reasoning)OpenAI
0.0%
$1.10 / $4.40
137 Solar Pro 3Upstage
0.0%
$0.15 / $0.60
137 Grok 4.3 (high reasoning)xAI
0.0%
$1.25 / $2.50
137 Grok 4.3 (no reasoning)xAI
0.0%
$1.25 / $2.50
137 MiMo-V2.5-Pro (reasoning on)Xiaomi
0.0%
$0.43 / $0.87
137 MiMo-V2.5Xiaomi
0.0%
$0.17 / $0.34

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of terminal tasks (software, system and data jobs) an agent completes, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not other harnesses or tools; one attempt per task.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.