Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of terminal tasks (software, system and data jobs) an agent completes, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
187 configurations of 112 models
Unit
% of tasks
Licence
Artificial Analysis commercial data licence

Terminal-Bench 4.0: Gemini 4 Argon

Top 15 of 112 results · % of tasks, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Claude Sonnet 5.5 (max reasoning)Anthropic 63.6%
  2. 2 Claude Opus 5.5 (max reasoning)Anthropic 59.6%
  3. 2 GPT-6 Astra (extra-high reasoning)OpenAI 59.6%
  4. 4 Gemini 4 Argon (high reasoning)Google 57.1%
  5. 5 GPT-6.1 Sol (max reasoning)OpenAI 56.1%
  6. 6 Claude Fable 5.1 (extra-high reasoning)Anthropic 55.1%
  7. 7 Claude Opus 5 (max reasoning)Anthropic 49.0%
  8. 8 GPT-6 Sol (max reasoning)OpenAI 43.9%
  9. 9 Claude Fable 5 (max reasoning)Anthropic 42.4%
  10. 10 GLM-5.3 (max reasoning)Z.ai 41.9%
  11. 11 GPT-5.6 Sol (max reasoning)OpenAI 39.9%
  12. 12 Qwen3.8-Max (0902)Alibaba 38.9%
  13. 13 GPT-5.6 Terra (max reasoning)OpenAI 35.4%
  14. 14 MiMo-V2.6-ProXiaomi 34.8%
  15. 15 Muse Spark 1.3 (max reasoning)Meta 33.3%

Full results

Artificial Analysis: Terminal-Bench 4.0, % of tasks, higher is better
#ModelTerminal-Bench 4.0
% of tasks, higher is better
Price
$ per million tokens, in / out
1 Claude Sonnet 5.5 (max reasoning)Anthropic · best of 5 settings
63.6%
$2 / $10
2 Claude Opus 5.5 (max reasoning)Anthropic · best of 5 settings
59.6%
$4 / $20
2 GPT-6 Astra (extra-high reasoning)OpenAI · best of 5 settings
59.6%
$10 / $50
4 Gemini 4 Argon (high reasoning)Google
57.1%
–
5 GPT-6.1 Sol (max reasoning)OpenAI · best of 5 settings
56.1%
$2 / $10
6 Claude Fable 5.1 (extra-high reasoning)Anthropic · best of 5 settings
55.1%
$10 / $50
7 Claude Opus 5 (max reasoning)Anthropic · best of 5 settings
49.0%
$5 / $25
8 GPT-6 Sol (max reasoning)OpenAI · best of 6 settings
43.9%
$2 / $10
9 Claude Fable 5 (max reasoning)Anthropic
42.4%
$10 / $50
10 GLM-5.3 (max reasoning)Z.ai · best of 2 settings
41.9%
$1.40 / $4.40
11 GPT-5.6 Sol (max reasoning)OpenAI · best of 5 settings
39.9%
$4 / $20
12 Qwen3.8-Max (0902)Alibaba
38.9%
$2 / $6
13 GPT-5.6 Terra (max reasoning)OpenAI · best of 6 settings
35.4%
$2 / $12
14 MiMo-V2.6-ProXiaomi
34.8%
$0.43 / $0.87
15 Muse Spark 1.3 (max reasoning)Meta · best of 2 settings
33.3%
$1.25 / $4.25
16 Claude Haiku 5.5 (max reasoning)Anthropic · best of 5 settings
32.8%
$0.10 / $0.50
16 GLM-5.3-FlashZ.ai
32.8%
$0.15 / $0.50
18 DeepSeek-V4.1-Flash (max reasoning)DeepSeek · best of 2 settings
26.8%
$0.30 / $1.20
18 Mistral Large 4Mistral AI
26.8%
$1.36 / $4.18
20 Grok 4.7 (extra-high reasoning)xAI · best of 3 settings
25.8%
$2 / $6
21 Qwen3.8-Flash-NextAlibaba
25.3%
–
22 MiMo-V2.6-FlashXiaomi
22.7%
$0.14 / $0.28
23 Claude Opus 4.8 (max reasoning)Anthropic
21.7%
$5 / $25
24 Grok 4.6 (high reasoning)xAI · best of 4 settings
21.2%
$2 / $6
25 Gemini 3.8 Flash (high reasoning)Google · best of 3 settings
19.7%
$1.50 / $7.50
26 Qwen3.8-Max (0803)Alibaba
18.7%
–
27 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek
14.6%
$1.42 / $2.83
27 GPT-5.5 (extra-high reasoning)OpenAI · best of 3 settings
14.6%
$5 / $30
29 Claude Sonnet 5 (max reasoning)Anthropic · best of 5 settings
14.1%
$2 / $10
29 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek · best of 2 settings
14.1%
$1.32 / $3.96
31 Gemini 3.7 Flash (high reasoning)Google
13.6%
$1.50 / $7.50
32 Kimi K3 (low reasoning)Moonshot AI · best of 2 settings
12.6%
$3 / $15
32 GPT-5.5 Instant (2026-06-26)OpenAI
12.6%
–
32 GPT-6 Luna (max reasoning)OpenAI · best of 6 settings
12.6%
$0.10 / $0.50
35 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
12.1%
$0.14 / $0.28
35 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
12.1%
$0.44 / $1.32
37 GPT-5.6 Luna (max reasoning)OpenAI · best of 6 settings
11.6%
$0.20 / $1.20
38 Qwen3.8-2.4T-A95BAlibaba
11.1%
$2 / $6
39 Grok 4.5 (high reasoning)xAI
10.6%
$2 / $6
40 Gemini 3.6 Flash (high reasoning)Google
7.1%
$1.50 / $7.50
40 Muse Spark 1.2 (extra-high reasoning)Meta
7.1%
$1.25 / $4.25
42 Gemini 3.5 Flash (high reasoning)Google
6.6%
$1.50 / $9
43 Muse Spark 1.1 (extra-high reasoning)Meta
6.1%
$1.25 / $4.25
44 Qwen3.8-27B (extra-high reasoning)Alibaba · best of 4 settings
5.6%
$0.50 / $3
45 Gemini 3.1 Pro PreviewGoogle
4.0%
$2 / $12
46 Claude Sonnet 4.6 (max reasoning)Anthropic
3.0%
$3 / $15
46 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek · best of 2 settings
3.0%
$0.14 / $0.28
48 MiniMax-M3MiniMax
2.0%
$0.30 / $1.20
48 GPT-5.4 mini (extra-high reasoning)OpenAI
2.0%
$0.75 / $4.50
48 GLM-5.1Z.ai
2.0%
$1.38 / $4.40
51 Qwen3.7-MaxAlibaba
1.5%
$1.48 / $4.42
52 Qwen3.7-PlusAlibaba
1.0%
$0.32 / $1.28
52 Gemini 3.5 Flash-LiteGoogle
1.0%
$0.30 / $2.50
52 Kimi K2.7 CodeMoonshot AI
1.0%
$0.95 / $4
52 Inkling SmallThinking Machines
1.0%
$0.45 / $1.20
52 Inkling (extra-high reasoning)Thinking Machines
1.0%
$0.95 / $4.05
52 GLM-5.2 (max reasoning)Z.ai
1.0%
$1.40 / $4.40
58 Qwen3.5-9BAlibaba
0.5%
$0.10 / $0.15
58 Trinity Large ThinkingArcee AI
0.5%
$0.25 / $0.80
58 Gemini 3.1 Flash-Lite PreviewGoogle
0.5%
$0.25 / $1.50
58 Muse Glimmer (high reasoning)Meta
0.5%
–
58 Kimi K2.6Moonshot AI
0.5%
$0.95 / $4
58 GPT-5.4 nano (extra-high reasoning)OpenAI
0.5%
$0.20 / $1.25
58 Hy3Tencent
0.5%
$0.14 / $0.58
58 Solar Pro 4Upstage
0.5%
$0.09 / $0.36
66 Qwen3-14B (reasoning on)Alibaba
0.0%
$0.12 / $0.24
66 Qwen3-235B-A22B-Thinking-2507Alibaba
0.0%
$0.30 / $3
66 Qwen3-30B-A3B-Thinking-2507Alibaba
0.0%
$0.20 / $2.40
66 Qwen3-32B (reasoning on)Alibaba
0.0%
$0.14 / $0.40
66 Qwen3.5-122B-A10BAlibaba
0.0%
$0.26 / $2.08
66 Qwen3.5-397B-A17BAlibaba
0.0%
$0.55 / $3.50
66 Qwen3.6-27BAlibaba
0.0%
$0.30 / $3.20
66 Qwen3.6-35B-A3BAlibaba
0.0%
$0.10 / $1
66 Qwen3-8B (reasoning on)Alibaba
0.0%
$0.12 / $0.46
66 Qwen3-Coder-NextAlibaba
0.0%
$0.18 / $0.90
66 Claude Haiku 4.5 (reasoning on)Anthropic
0.0%
$1 / $5
66 Claude Sonnet 4.5 (reasoning on)Anthropic
0.0%
$3 / $15
66 DeepSeek-V3-0324DeepSeek
0.0%
$0.25 / $1
66 DeepSeek-V3DeepSeek
0.0%
$0.26 / $1.03
66 DeepSeek-R1DeepSeek
0.0%
$0.70 / $2.50
66 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
0.0%
$0.27 / $1
66 Gemini 2.5 ProGoogle
0.0%
$1.25 / $10
66 Gemma 3 12BGoogle
0.0%
$0.05 / $0.15
66 Gemma 3 27BGoogle
0.0%
$0.12 / $0.20
66 Gemma 4 31BGoogle
0.0%
$0.14 / $0.40
66 Gemma 4 E4BGoogle
0.0%
–
66 Granite 4.2 8BIBM
0.0%
$0.06 / $0.25
66 Mercury 2Inception
0.0%
$0.25 / $0.75
66 Llama 4 MaverickMeta
0.0%
$0.27 / $0.85
66 Llama 4 ScoutMeta
0.0%
$0.18 / $0.59
66 MiniMax-M2.7MiniMax
0.0%
$0.30 / $1.20
66 Devstral 2Mistral AI
0.0%
$0.40 / $2
66 Devstral Small 2Mistral AI
0.0%
–
66 Ministral 3 14BMistral AI
0.0%
$0.20 / $0.20
66 Ministral 3 3BMistral AI
0.0%
$0.10 / $0.10
66 Ministral 3 8BMistral AI
0.0%
$0.15 / $0.15
66 Mistral Large 3Mistral AI
0.0%
$0.50 / $1.50
66 Mistral Medium 3.1Mistral AI
0.0%
$0.40 / $2
66 Mistral Medium 3.5Mistral AI
0.0%
$1.50 / $7.50
66 Mistral Small 4Mistral AI
0.0%
$0.15 / $0.60
66 Mistral Small 3.1 24BMistral AI
0.0%
$0.35 / $0.56
66 Mistral Small 3.2 24BMistral AI
0.0%
$0.094 / $0.25
66 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA
0.0%
$0.05 / $0.20
66 Nemotron 3 Super 120B A12BNVIDIA
0.0%
$0.085 / $0.40
66 GPT-5 mini (high reasoning)OpenAI
0.0%
$0.25 / $2
66 gpt-oss-120b (high reasoning)OpenAI
0.0%
$0.15 / $0.60
66 gpt-oss-20b (high reasoning)OpenAI
0.0%
$0.03 / $0.15
66 o3-mini (high reasoning)OpenAI
0.0%
$1.10 / $4.40
66 Solar Pro 3Upstage
0.0%
$0.15 / $0.60
66 Grok 4.3 (high reasoning)xAI · best of 2 settings
0.0%
$1.25 / $2.50
66 MiMo-V2.5-Pro (reasoning on)Xiaomi
0.0%
$0.43 / $0.87
66 MiMo-V2.5Xiaomi
0.0%
$0.17 / $0.34

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of terminal tasks (software, system and data jobs) an agent completes, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not other harnesses or tools; one attempt per task.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.