Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of scientific coding sub-problems solved, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
191 configurations of 113 models
Unit
% of problems
Licence
Artificial Analysis commercial data licence

SciCode: Claude Haiku 5.5

Top 15 of 113 results · % of problems, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Claude Opus 5.5 (max reasoning)Anthropic 66.9%
  2. 2 Claude Fable 5.1 (max reasoning)Anthropic 63.1%
  3. 3 Gemini 4 Argon (high reasoning)Google 61.8%
  4. 4 Claude Fable 5 (max reasoning)Anthropic 61.0%
  5. 4 Claude Sonnet 5.5 (max reasoning)Anthropic 61.0%
  6. 6 MiMo-V2.6-ProXiaomi 60.9%
  7. 7 Gemini 3.7 Flash (medium reasoning)Google 59.8%
  8. 8 Muse Spark 1.3 (extra-high reasoning)Meta 59.7%
  9. 9 Kimi K3 (max reasoning)Moonshot AI 59.5%
  10. 10 GLM-5.3 (max reasoning)Z.ai 59.0%
  11. 11 Muse Spark 1.1 (extra-high reasoning)Meta 58.8%
  12. 12 Gemini 3.1 Pro PreviewGoogle 58.7%
  13. 13 GPT-5.6 Sol (high reasoning)OpenAI 57.8%
  14. 13 Grok 4.7 (high reasoning)xAI 57.8%
  15. 15 GPT-6 Sol (max reasoning)OpenAI 57.6%
  16. 23 Claude Haiku 5.5 (max reasoning)Anthropic 55.0%

Full results

Artificial Analysis: SciCode, % of problems, higher is better
#ModelSciCode
% of problems, higher is better
Price
$ per million tokens, in / out
1 Claude Opus 5.5 (max reasoning)Anthropic · best of 5 settings
66.9%
$4 / $20
2 Claude Fable 5.1 (max reasoning)Anthropic · best of 5 settings
63.1%
$10 / $50
3 Gemini 4 Argon (high reasoning)Google
61.8%
–
4 Claude Fable 5 (max reasoning)Anthropic
61.0%
$10 / $50
4 Claude Sonnet 5.5 (max reasoning)Anthropic · best of 5 settings
61.0%
$2 / $10
6 MiMo-V2.6-ProXiaomi
60.9%
$0.43 / $0.87
7 Gemini 3.7 Flash (medium reasoning)Google · best of 3 settings
59.8%
$1.50 / $7.50
8 Muse Spark 1.3 (extra-high reasoning)Meta · best of 2 settings
59.7%
$1.25 / $4.25
9 Kimi K3 (max reasoning)Moonshot AI · best of 2 settings
59.5%
$3 / $15
10 GLM-5.3 (max reasoning)Z.ai · best of 2 settings
59.0%
$1.40 / $4.40
11 Muse Spark 1.1 (extra-high reasoning)Meta
58.8%
$1.25 / $4.25
12 Gemini 3.1 Pro PreviewGoogle
58.7%
$2 / $12
13 GPT-5.6 Sol (high reasoning)OpenAI · best of 6 settings
57.8%
$4 / $20
13 Grok 4.7 (high reasoning)xAI · best of 3 settings
57.8%
$2 / $6
15 GPT-6 Sol (max reasoning)OpenAI · best of 6 settings
57.6%
$2 / $10
16 Muse Spark 1.2 (extra-high reasoning)Meta
57.4%
$1.25 / $4.25
17 Gemini 3.8 Flash (high reasoning)Google · best of 3 settings
56.6%
$1.50 / $7.50
18 GPT-6 Astra (max reasoning)OpenAI · best of 5 settings
56.5%
$10 / $50
18 Grok 4.6 (high reasoning)xAI · best of 4 settings
56.5%
$2 / $6
20 Claude Opus 5 (max reasoning)Anthropic · best of 5 settings
56.4%
$5 / $25
21 GPT-5.5 (high reasoning)OpenAI · best of 3 settings
56.1%
$5 / $30
22 GPT-6.1 Sol (high reasoning)OpenAI · best of 5 settings
55.8%
$2 / $10
23 Claude Haiku 5.5 (max reasoning)Anthropic · best of 5 settings
55.0%
$0.10 / $0.50
23 GPT-5.6 Terra (max reasoning)OpenAI · best of 6 settings
55.0%
$2 / $12
23 Grok 4.5 (high reasoning)xAI
55.0%
$2 / $6
26 GPT-6 Luna (max reasoning)OpenAI · best of 6 settings
54.6%
$0.10 / $0.50
27 Claude Opus 4.8 (max reasoning)Anthropic
54.4%
$5 / $25
28 Claude Sonnet 5 (high reasoning)Anthropic · best of 5 settings
54.3%
$2 / $10
29 Mistral Large 4Mistral AI
54.2%
$1.36 / $4.18
30 Qwen3.8-2.4T-A95BAlibaba
54.1%
$2 / $6
31 Gemini 3.5 Flash (high reasoning)Google
53.9%
$1.50 / $9
32 GPT-5.6 Luna (max reasoning)OpenAI · best of 6 settings
53.6%
$0.20 / $1.20
33 Gemini 3.6 Flash (high reasoning)Google
53.4%
$1.50 / $7.50
34 Qwen3.8-Max (0803)Alibaba
53.2%
–
35 GPT-5.5 Instant (2026-06-26)OpenAI
52.5%
–
36 Qwen3.8-Max (0902)Alibaba
52.1%
$2 / $6
36 GPT-5.4 mini (extra-high reasoning)OpenAI
52.1%
$0.75 / $4.50
38 DeepSeek-V4.1-Flash (max reasoning)DeepSeek · best of 2 settings
51.9%
$0.30 / $1.20
39 GLM-5.3-FlashZ.ai
51.6%
$0.15 / $0.50
40 Kimi K2.6Moonshot AI
51.5%
$0.95 / $4
41 MiMo-V2.6-FlashXiaomi
51.3%
$0.14 / $0.28
42 GLM-5.2 (max reasoning)Z.ai
51.2%
$1.40 / $4.40
43 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek · best of 2 settings
51.0%
$1.32 / $3.96
44 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek
50.8%
$1.42 / $2.83
45 Qwen3.8-Flash-NextAlibaba
50.6%
–
45 MiMo-V2.5-Pro (reasoning on)Xiaomi
50.6%
$0.43 / $0.87
47 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
50.3%
$0.14 / $0.28
48 Claude Sonnet 4.6 (max reasoning)Anthropic
50.1%
$3 / $15
48 MiniMax-M2.7MiniMax
50.1%
$0.30 / $1.20
50 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
49.7%
$0.44 / $1.32
50 Inkling SmallThinking Machines
49.7%
$0.45 / $1.20
52 Qwen3.7-MaxAlibaba
49.5%
$1.48 / $4.42
53 Hy3Tencent
48.6%
$0.14 / $0.58
54 Grok 4.3 (high reasoning)xAI · best of 2 settings
48.3%
$1.25 / $2.50
55 Kimi K2.7 CodeMoonshot AI
47.8%
$0.95 / $4
56 GPT-5.4 nano (extra-high reasoning)OpenAI
47.2%
$0.20 / $1.25
57 MiniMax-M3MiniMax
47.1%
$0.30 / $1.20
58 Inkling (extra-high reasoning)Thinking Machines
47.0%
$0.95 / $4.05
59 Qwen3.8-27B (extra-high reasoning)Alibaba · best of 4 settings
46.6%
$0.50 / $3
60 Gemini 2.5 ProGoogle
46.3%
$1.25 / $10
61 Qwen3.7-PlusAlibaba
46.1%
$0.32 / $1.28
62 Claude Sonnet 4.5 (reasoning on)Anthropic
45.7%
$3 / $15
63 Gemma 4 31BGoogle
45.5%
$0.14 / $0.40
64 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek · best of 2 settings
45.3%
$0.14 / $0.28
65 Muse Glimmer (high reasoning)Meta
44.9%
–
66 Qwen3.5-397B-A17BAlibaba
44.8%
$0.55 / $3.50
66 GLM-5.1Z.ai
44.8%
$1.38 / $4.40
68 Solar Pro 4Upstage
44.6%
$0.09 / $0.36
69 MiMo-V2.5Xiaomi
43.9%
$0.17 / $0.34
70 Gemini 3.1 Flash-Lite PreviewGoogle
43.4%
$0.25 / $1.50
71 Qwen3.6-27BAlibaba
42.8%
$0.30 / $3.20
71 o3-mini (high reasoning)OpenAI
42.8%
$1.10 / $4.40
73 Claude Haiku 4.5 (reasoning on)Anthropic
42.2%
$1 / $5
74 Qwen3-235B-A22B-Thinking-2507Alibaba
41.4%
$0.30 / $3
75 Gemini 3.5 Flash-LiteGoogle
41.3%
$0.30 / $2.50
76 Trinity Large ThinkingArcee AI
40.6%
$0.25 / $0.80
77 Mistral Medium 3.5Mistral AI
40.2%
$1.50 / $7.50
78 Gemma 4 26B A4BGoogle
40.0%
$0.10 / $0.30
79 Qwen3.5-122B-A10BAlibaba
39.7%
$0.26 / $2.08
80 DeepSeek-V3-0324DeepSeek
39.0%
$0.25 / $1
80 GPT-5 mini (high reasoning)OpenAI
39.0%
$0.25 / $2
82 gpt-oss-20b (high reasoning)OpenAI
38.9%
$0.03 / $0.15
83 Mistral Small 4Mistral AI
38.8%
$0.15 / $0.60
84 DeepSeek-R1DeepSeek
38.3%
$0.70 / $2.50
85 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
38.0%
$0.27 / $1
86 Mercury 2Inception
37.7%
$0.25 / $0.75
87 Qwen3.6-35B-A3BAlibaba
36.6%
$0.10 / $1
87 Mistral Large 3Mistral AI
36.6%
$0.50 / $1.50
89 Qwen3-Coder-NextAlibaba
36.2%
$0.18 / $0.90
89 Nemotron 3 Super 120B A12BNVIDIA
36.2%
$0.085 / $0.40
91 Qwen3-32B (reasoning on)Alibaba
36.0%
$0.14 / $0.40
92 DeepSeek-V3DeepSeek
35.8%
$0.26 / $1.03
93 gpt-oss-120b (high reasoning)OpenAI
34.0%
$0.15 / $0.60
94 Qwen3-30B-A3B-Thinking-2507Alibaba
33.0%
$0.20 / $2.40
95 Devstral 2Mistral AI
32.8%
$0.40 / $2
96 Devstral Small 2Mistral AI
32.4%
–
96 Mistral Medium 3.1Mistral AI
32.4%
$0.40 / $2
98 Llama 4 MaverickMeta
31.7%
$0.27 / $0.85
99 Granite 4.2 8BIBM
31.5%
$0.06 / $0.25
100 Qwen3-14B (reasoning on)Alibaba
30.7%
$0.12 / $0.24
101 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA
30.6%
$0.05 / $0.20
102 Qwen3.5-9BAlibaba
29.5%
$0.10 / $0.15
103 Mistral Small 3.2 24BMistral AI
28.6%
$0.094 / $0.25
104 Qwen3-8B (reasoning on)Alibaba
27.9%
$0.12 / $0.46
105 Mistral Small 3.1 24BMistral AI
27.8%
$0.35 / $0.56
106 Solar Pro 3Upstage
25.5%
$0.15 / $0.60
107 Gemma 4 E4BGoogle
24.4%
–
108 Ministral 3 14BMistral AI
23.8%
$0.20 / $0.20
109 Gemma 3 27BGoogle
23.3%
$0.12 / $0.20
110 Llama 4 ScoutMeta
21.3%
$0.18 / $0.59
111 Ministral 3 8BMistral AI
20.7%
$0.15 / $0.15
112 Gemma 3 12BGoogle
16.4%
$0.05 / $0.15
113 Ministral 3 3BMistral AI
15.3%
$0.10 / $0.10

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model is shown at its best setting; show every setting. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of scientific coding sub-problems solved, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not general software engineering.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.