Benchmarks / Artificial Analysis

Reported by Artificial Analysis

Artificial Analysis

Share of scientific coding sub-problems solved, as run by Artificial Analysis with its own harness and prompts.

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Results
191 configurations of 113 models
Unit
% of problems
Licence
Artificial Analysis commercial data licence

SciCode: GPT-5.6 Sol

Top 16 of 191 results · % of problems, higher is better. Choose a model to highlight it.Clear highlight

  1. 1 Claude Opus 5.5 (max reasoning)Anthropic 66.9%
  2. 2 Claude Opus 5.5 (extra-high reasoning)Anthropic 65.0%
  3. 3 Claude Fable 5.1 (max reasoning)Anthropic 63.1%
  4. 4 Gemini 4 Argon (high reasoning)Google 61.8%
  5. 5 Claude Fable 5 (max reasoning)Anthropic 61.0%
  6. 5 Claude Sonnet 5.5 (max reasoning)Anthropic 61.0%
  7. 7 Claude Fable 5.1 (extra-high reasoning)Anthropic 60.9%
  8. 7 MiMo-V2.6-ProXiaomi 60.9%
  9. 9 Claude Opus 5.5 (high reasoning)Anthropic 60.4%
  10. 10 Gemini 3.7 Flash (medium reasoning)Google 59.8%
  11. 11 Muse Spark 1.3 (extra-high reasoning)Meta 59.7%
  12. 12 Kimi K3 (max reasoning)Moonshot AI 59.5%
  13. 13 Claude Opus 5.5 (medium reasoning)Anthropic 59.3%
  14. 14 GLM-5.3 (max reasoning)Z.ai 59.0%
  15. 15 Muse Spark 1.1 (extra-high reasoning)Meta 58.8%
  16. 20 GPT-5.6 Sol (high reasoning)OpenAI 57.8%
  17. 23 GPT-5.6 Sol (medium reasoning)OpenAI 57.4%
  18. 28 GPT-5.6 Sol (max reasoning)OpenAI 57.1%
  19. 28 GPT-5.6 Sol (extra-high reasoning)OpenAI 57.1%
  20. 34 GPT-5.6 Sol (low reasoning)OpenAI 56.4%
  21. 118 GPT-5.6 Sol (no reasoning)OpenAI 47.7%

Full results

Artificial Analysis: SciCode, % of problems, higher is better
#ModelSciCode
% of problems, higher is better
Price
$ per million tokens, in / out
1 Claude Opus 5.5 (max reasoning)Anthropic
66.9%
$4 / $20
2 Claude Opus 5.5 (extra-high reasoning)Anthropic
65.0%
$4 / $20
3 Claude Fable 5.1 (max reasoning)Anthropic
63.1%
$10 / $50
4 Gemini 4 Argon (high reasoning)Google
61.8%
–
5 Claude Fable 5 (max reasoning)Anthropic
61.0%
$10 / $50
5 Claude Sonnet 5.5 (max reasoning)Anthropic
61.0%
$2 / $10
7 Claude Fable 5.1 (extra-high reasoning)Anthropic
60.9%
$10 / $50
7 MiMo-V2.6-ProXiaomi
60.9%
$0.43 / $0.87
9 Claude Opus 5.5 (high reasoning)Anthropic
60.4%
$4 / $20
10 Gemini 3.7 Flash (medium reasoning)Google
59.8%
$1.50 / $7.50
11 Muse Spark 1.3 (extra-high reasoning)Meta
59.7%
$1.25 / $4.25
12 Kimi K3 (max reasoning)Moonshot AI
59.5%
$3 / $15
13 Claude Opus 5.5 (medium reasoning)Anthropic
59.3%
$4 / $20
14 GLM-5.3 (max reasoning)Z.ai
59.0%
$1.40 / $4.40
15 Muse Spark 1.1 (extra-high reasoning)Meta
58.8%
$1.25 / $4.25
15 Muse Spark 1.3 (max reasoning)Meta
58.8%
$1.25 / $4.25
17 Claude Fable 5.1 (high reasoning)Anthropic
58.7%
$10 / $50
17 Gemini 3.1 Pro PreviewGoogle
58.7%
$2 / $12
19 Claude Opus 5.5 (low reasoning)Anthropic
58.6%
$4 / $20
20 GPT-5.6 Sol (high reasoning)OpenAI
57.8%
$4 / $20
20 Grok 4.7 (high reasoning)xAI
57.8%
$2 / $6
22 GPT-6 Sol (max reasoning)OpenAI
57.6%
$2 / $10
23 Muse Spark 1.2 (extra-high reasoning)Meta
57.4%
$1.25 / $4.25
23 GPT-5.6 Sol (medium reasoning)OpenAI
57.4%
$4 / $20
23 Grok 4.7 (extra-high reasoning)xAI
57.4%
$2 / $6
26 Claude Sonnet 5.5 (extra-high reasoning)Anthropic
57.3%
$2 / $10
27 Gemini 3.7 Flash (high reasoning)Google
57.2%
$1.50 / $7.50
28 GPT-5.6 Sol (max reasoning)OpenAI
57.1%
$4 / $20
28 GPT-5.6 Sol (extra-high reasoning)OpenAI
57.1%
$4 / $20
30 Claude Fable 5.1 (low reasoning)Anthropic
56.7%
$10 / $50
31 Gemini 3.8 Flash (high reasoning)Google
56.6%
$1.50 / $7.50
32 GPT-6 Astra (max reasoning)OpenAI
56.5%
$10 / $50
32 Grok 4.6 (high reasoning)xAI
56.5%
$2 / $6
34 Claude Fable 5.1 (medium reasoning)Anthropic
56.4%
$10 / $50
34 Claude Opus 5 (max reasoning)Anthropic
56.4%
$5 / $25
34 GPT-5.6 Sol (low reasoning)OpenAI
56.4%
$4 / $20
37 GPT-5.5 (high reasoning)OpenAI
56.1%
$5 / $30
38 Grok 4.6 (medium reasoning)xAI
55.9%
$2 / $6
39 GPT-5.5 (extra-high reasoning)OpenAI
55.8%
$5 / $30
39 GPT-6.1 Sol (high reasoning)OpenAI
55.8%
$2 / $10
41 Claude Opus 5 (extra-high reasoning)Anthropic
55.7%
$5 / $25
41 Gemini 3.7 Flash (low reasoning)Google
55.7%
$1.50 / $7.50
41 GPT-6.1 Sol (extra-high reasoning)OpenAI
55.7%
$2 / $10
41 GPT-6 Astra (extra-high reasoning)OpenAI
55.7%
$10 / $50
45 Claude Opus 5 (high reasoning)Anthropic
55.4%
$5 / $25
45 GPT-6 Astra (high reasoning)OpenAI
55.4%
$10 / $50
47 Gemini 3.8 Flash (medium reasoning)Google
55.1%
$1.50 / $7.50
47 GPT-6 Sol (extra-high reasoning)OpenAI
55.1%
$2 / $10
49 Claude Haiku 5.5 (max reasoning)Anthropic
55.0%
$0.10 / $0.50
49 Gemini 3.8 Flash (low reasoning)Google
55.0%
$1.50 / $7.50
49 GPT-5.6 Terra (max reasoning)OpenAI
55.0%
$2 / $12
49 Grok 4.5 (high reasoning)xAI
55.0%
$2 / $6
53 GPT-6 Sol (high reasoning)OpenAI
54.9%
$2 / $10
53 Grok 4.7 (low reasoning)xAI
54.9%
$2 / $6
55 GPT-6 Luna (max reasoning)OpenAI
54.6%
$0.10 / $0.50
56 GPT-5.5 (medium reasoning)OpenAI
54.5%
$5 / $30
57 Claude Opus 4.8 (max reasoning)Anthropic
54.4%
$5 / $25
58 Claude Sonnet 5 (high reasoning)Anthropic
54.3%
$2 / $10
58 Claude Sonnet 5 (max reasoning)Anthropic
54.3%
$2 / $10
60 Mistral Large 4Mistral AI
54.2%
$1.36 / $4.18
60 GPT-6.1 Sol (max reasoning)OpenAI
54.2%
$2 / $10
60 GPT-6 Astra (medium reasoning)OpenAI
54.2%
$10 / $50
63 Qwen3.8-2.4T-A95BAlibaba
54.1%
$2 / $6
63 Claude Sonnet 5 (extra-high reasoning)Anthropic
54.1%
$2 / $10
63 GPT-6 Astra (low reasoning)OpenAI
54.1%
$10 / $50
66 Gemini 3.5 Flash (high reasoning)Google
53.9%
$1.50 / $9
67 GPT-6 Sol (medium reasoning)OpenAI
53.8%
$2 / $10
68 Claude Sonnet 5.5 (high reasoning)Anthropic
53.7%
$2 / $10
69 GPT-5.6 Luna (max reasoning)OpenAI
53.6%
$0.20 / $1.20
70 Gemini 3.6 Flash (high reasoning)Google
53.4%
$1.50 / $7.50
71 Qwen3.8-Max (0803)Alibaba
53.2%
–
71 GPT-6.1 Sol (low reasoning)OpenAI
53.2%
$2 / $10
71 GPT-6.1 Sol (medium reasoning)OpenAI
53.2%
$2 / $10
74 Grok 4.6 (extra-high reasoning)xAI
53.0%
$2 / $6
75 Claude Sonnet 5.5 (medium reasoning)Anthropic
52.9%
$2 / $10
76 Kimi K3 (low reasoning)Moonshot AI
52.7%
$3 / $15
77 GPT-5.5 Instant (2026-06-26)OpenAI
52.5%
–
78 GPT-5.6 Terra (high reasoning)OpenAI
52.4%
$2 / $12
79 GPT-5.6 Terra (extra-high reasoning)OpenAI
52.3%
$2 / $12
80 Qwen3.8-Max (0902)Alibaba
52.1%
$2 / $6
80 GPT-5.4 mini (extra-high reasoning)OpenAI
52.1%
$0.75 / $4.50
82 DeepSeek-V4.1-Flash (max reasoning)DeepSeek
51.9%
$0.30 / $1.20
83 Claude Haiku 5.5 (extra-high reasoning)Anthropic
51.7%
$0.10 / $0.50
83 GPT-6 Luna (extra-high reasoning)OpenAI
51.7%
$0.10 / $0.50
85 Claude Sonnet 5 (medium reasoning)Anthropic
51.6%
$2 / $10
85 GPT-5.6 Luna (high reasoning)OpenAI
51.6%
$0.20 / $1.20
85 GLM-5.3-FlashZ.ai
51.6%
$0.15 / $0.50
88 Claude Opus 5 (medium reasoning)Anthropic
51.5%
$5 / $25
88 Kimi K2.6Moonshot AI
51.5%
$0.95 / $4
90 MiMo-V2.6-FlashXiaomi
51.3%
$0.14 / $0.28
91 GLM-5.2 (max reasoning)Z.ai
51.2%
$1.40 / $4.40
92 DeepSeek-V4-Pro (0813, max reasoning)DeepSeek
51.0%
$1.32 / $3.96
93 GPT-6 Luna (medium reasoning)OpenAI
50.9%
$0.10 / $0.50
94 DeepSeek-V4-Pro (0423, max reasoning)DeepSeek
50.8%
$1.42 / $2.83
95 Qwen3.8-Flash-NextAlibaba
50.6%
–
95 MiMo-V2.5-Pro (reasoning on)Xiaomi
50.6%
$0.43 / $0.87
97 GPT-5.6 Luna (extra-high reasoning)OpenAI
50.5%
$0.20 / $1.20
97 GPT-5.6 Terra (medium reasoning)OpenAI
50.5%
$2 / $12
99 DeepSeek-V4-Flash (0731, max reasoning)DeepSeek
50.3%
$0.14 / $0.28
99 GPT-6 Luna (high reasoning)OpenAI
50.3%
$0.10 / $0.50
101 GPT-6 Sol (low reasoning)OpenAI
50.2%
$2 / $10
102 Claude Sonnet 4.6 (max reasoning)Anthropic
50.1%
$3 / $15
102 Claude Sonnet 5 (low reasoning)Anthropic
50.1%
$2 / $10
102 MiniMax-M2.7MiniMax
50.1%
$0.30 / $1.20
105 GPT-5.6 Terra (low reasoning)OpenAI
49.9%
$2 / $12
106 DeepSeek-V4-Flash-Vision-Exp (max reasoning)DeepSeek
49.7%
$0.44 / $1.32
106 Inkling SmallThinking Machines
49.7%
$0.45 / $1.20
108 Qwen3.7-MaxAlibaba
49.5%
$1.48 / $4.42
109 Grok 4.6 (low reasoning)xAI
49.4%
$2 / $6
110 Claude Haiku 5.5 (low reasoning)Anthropic
49.2%
$0.10 / $0.50
110 Claude Opus 5 (low reasoning)Anthropic
49.2%
$5 / $25
112 Claude Sonnet 5.5 (low reasoning)Anthropic
49.1%
$2 / $10
113 Claude Haiku 5.5 (medium reasoning)Anthropic
49.0%
$0.10 / $0.50
114 Claude Haiku 5.5 (high reasoning)Anthropic
48.7%
$0.10 / $0.50
115 Hy3Tencent
48.6%
$0.14 / $0.58
116 Grok 4.3 (high reasoning)xAI
48.3%
$1.25 / $2.50
117 Kimi K2.7 CodeMoonshot AI
47.8%
$0.95 / $4
118 GPT-5.6 Sol (no reasoning)OpenAI
47.7%
$4 / $20
119 GPT-6 Sol (no reasoning)OpenAI
47.3%
$2 / $10
120 GPT-5.4 nano (extra-high reasoning)OpenAI
47.2%
$0.20 / $1.25
121 MiniMax-M3MiniMax
47.1%
$0.30 / $1.20
122 Inkling (extra-high reasoning)Thinking Machines
47.0%
$0.95 / $4.05
123 GPT-6 Luna (low reasoning)OpenAI
46.9%
$0.10 / $0.50
124 GPT-5.6 Luna (medium reasoning)OpenAI
46.8%
$0.20 / $1.20
125 Qwen3.8-27B (extra-high reasoning)Alibaba
46.6%
$0.50 / $3
126 Gemini 2.5 ProGoogle
46.3%
$1.25 / $10
127 Qwen3.7-PlusAlibaba
46.1%
$0.32 / $1.28
127 GPT-5.6 Luna (low reasoning)OpenAI
46.1%
$0.20 / $1.20
129 Claude Sonnet 4.5 (reasoning on)Anthropic
45.7%
$3 / $15
130 Gemma 4 31BGoogle
45.5%
$0.14 / $0.40
131 DeepSeek-V4-Flash (0423, max reasoning)DeepSeek
45.3%
$0.14 / $0.28
132 GPT-5.6 Terra (no reasoning)OpenAI
45.1%
$2 / $12
133 Muse Glimmer (high reasoning)Meta
44.9%
–
134 Qwen3.5-397B-A17BAlibaba
44.8%
$0.55 / $3.50
134 GLM-5.1Z.ai
44.8%
$1.38 / $4.40
136 Solar Pro 4Upstage
44.6%
$0.09 / $0.36
137 MiMo-V2.5Xiaomi
43.9%
$0.17 / $0.34
138 Gemini 3.1 Flash-Lite PreviewGoogle
43.4%
$0.25 / $1.50
139 GPT-6 Luna (no reasoning)OpenAI
43.1%
$0.10 / $0.50
140 Qwen3.6-27BAlibaba
42.8%
$0.30 / $3.20
140 o3-mini (high reasoning)OpenAI
42.8%
$1.10 / $4.40
142 Claude Haiku 4.5 (reasoning on)Anthropic
42.2%
$1 / $5
143 GLM-5.3 (low reasoning)Z.ai
42.0%
$1.40 / $4.40
144 Qwen3-235B-A22B-Thinking-2507Alibaba
41.4%
$0.30 / $3
145 Gemini 3.5 Flash-LiteGoogle
41.3%
$0.30 / $2.50
146 Trinity Large ThinkingArcee AI
40.6%
$0.25 / $0.80
147 GPT-5.6 Luna (no reasoning)OpenAI
40.4%
$0.20 / $1.20
148 DeepSeek-V4-Flash (0423, high reasoning)DeepSeek
40.2%
$0.14 / $0.28
148 Mistral Medium 3.5Mistral AI
40.2%
$1.50 / $7.50
150 Qwen3.8-27B (low reasoning)Alibaba
40.0%
$0.50 / $3
150 Gemma 4 26B A4BGoogle
40.0%
$0.10 / $0.30
152 DeepSeek-V4-Pro (0813, no reasoning)DeepSeek
39.9%
$1.32 / $3.96
153 Qwen3.5-122B-A10BAlibaba
39.7%
$0.26 / $2.08
154 Grok 4.3 (no reasoning)xAI
39.4%
$1.25 / $2.50
155 Qwen3.8-27B (medium reasoning)Alibaba
39.0%
$0.50 / $3
155 DeepSeek-V3-0324DeepSeek
39.0%
$0.25 / $1
155 GPT-5 mini (high reasoning)OpenAI
39.0%
$0.25 / $2
158 gpt-oss-20b (high reasoning)OpenAI
38.9%
$0.03 / $0.15
159 Mistral Small 4Mistral AI
38.8%
$0.15 / $0.60
160 DeepSeek-R1DeepSeek
38.3%
$0.70 / $2.50
161 DeepSeek-V3.1-Terminus (reasoning on)DeepSeek
38.0%
$0.27 / $1
162 Mercury 2Inception
37.7%
$0.25 / $0.75
163 Qwen3.6-35B-A3BAlibaba
36.6%
$0.10 / $1
163 Mistral Large 3Mistral AI
36.6%
$0.50 / $1.50
165 Qwen3.8-27B (no reasoning)Alibaba
36.2%
$0.50 / $3
165 Qwen3-Coder-NextAlibaba
36.2%
$0.18 / $0.90
165 Nemotron 3 Super 120B A12BNVIDIA
36.2%
$0.085 / $0.40
168 Qwen3-32B (reasoning on)Alibaba
36.0%
$0.14 / $0.40
169 DeepSeek-V3DeepSeek
35.8%
$0.26 / $1.03
170 DeepSeek-V4.1-Flash (no reasoning)DeepSeek
35.5%
$0.30 / $1.20
171 gpt-oss-120b (high reasoning)OpenAI
34.0%
$0.15 / $0.60
172 Qwen3-30B-A3B-Thinking-2507Alibaba
33.0%
$0.20 / $2.40
173 Devstral 2Mistral AI
32.8%
$0.40 / $2
174 Devstral Small 2Mistral AI
32.4%
–
174 Mistral Medium 3.1Mistral AI
32.4%
$0.40 / $2
176 Llama 4 MaverickMeta
31.7%
$0.27 / $0.85
177 Granite 4.2 8BIBM
31.5%
$0.06 / $0.25
178 Qwen3-14B (reasoning on)Alibaba
30.7%
$0.12 / $0.24
179 Nemotron 3 Nano 30B A3B (reasoning on)NVIDIA
30.6%
$0.05 / $0.20
180 Qwen3.5-9BAlibaba
29.5%
$0.10 / $0.15
181 Mistral Small 3.2 24BMistral AI
28.6%
$0.094 / $0.25
182 Qwen3-8B (reasoning on)Alibaba
27.9%
$0.12 / $0.46
183 Mistral Small 3.1 24BMistral AI
27.8%
$0.35 / $0.56
184 Solar Pro 3Upstage
25.5%
$0.15 / $0.60
185 Gemma 4 E4BGoogle
24.4%
–
186 Ministral 3 14BMistral AI
23.8%
$0.20 / $0.20
187 Gemma 3 27BGoogle
23.3%
$0.12 / $0.20
188 Llama 4 ScoutMeta
21.3%
$0.18 / $0.59
189 Ministral 3 8BMistral AI
20.7%
$0.15 / $0.15
190 Gemma 3 12BGoogle
16.4%
$0.05 / $0.15
191 Ministral 3 3BMistral AI
15.3%
$0.10 / $0.10

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Results as published by Artificial Analysis; we do not re-run them.

What it measures

Share of scientific coding sub-problems solved, as run by Artificial Analysis with its own harness and prompts.

What it does not measure

Not general software engineering.

282 results from Artificial Analysis not ranked here · show why

We rank a result only when we can tie it to a specific model you can use. These are left out:

  • Not on sale through the API providers we track: 268
  • A different snapshot or variant from the model we list: 13
  • An unusual combination of settings: 1

Data sourced from Artificial Analysis. Licence: Artificial Analysis commercial data licence.