Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

SciCode benchmark: model scores and methodology

A scientific-programming benchmark requiring Python solutions to research-oriented computational problems.

Current published leader
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Top score
60%
Primary metric
Pass rate · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current SciCode model scores

250 configurations · 182 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 250 current configurations with a reported SciCode score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on SciCode

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. These results sit in a narrow band, so the scale below is zoomed — read each mark against the labelled axis, not the left edge. Marks are placed on the full source precision, so two models sharing a rounded label can still sit at different points. Score labels reproduce Artificial Analysis's rounded display value.

  1. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 60%
  2. Gemini 3.1 Pro (Preview) 59%
  3. Kimi K3 59%
  4. Muse Spark 1.1 xhigh 58%
  5. GPT-5.6 Sol high 57%
  6. Claude Opus 5 56%
  7. Grok 4.5 54%
  8. GPT-5.6 Terra 54%
  9. Claude Sonnet 5 Adaptive Reasoning, Max Effort 54%
  10. GPT-5.3 Codex xhigh 53%
  11. Gemini 3.5 Flash high 53%
  12. Gemini 3.6 Flash high 53%
  13. GPT-5.6 Luna 53%
  14. Muse Spark 52%
  15. GLM-5.2 max 50%
  16. MiMo-V2.5-Pro 50%
  17. DeepSeek V4 Pro Reasoning, Max Effort 50%
  18. Qwen3.7 Max 49%
  19. GPT-5.5 Instant June 2026 49%
  20. Hy3 48%
43.2 52.4 61.7

Pass rate · higher is better · zoomed scale

Current configurations with a reported SciCode value; missing values are omitted.

How to interpret the result

What do SciCode results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher pass rate is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read SciCode as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationPass rate
1 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Leader Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 60%
2 Gemini 3.1 Pro (Preview) Published configuration 59%
3 Kimi K3 Published configuration 59%
4 Muse Spark 1.1 (xhigh) xhigh 58%
5 GPT-5.6 Sol (high) high 57%
6 GPT-5.6 Sol (medium) medium 56%
7 GPT-5.6 Sol max 56%
8 GPT-5.6 Sol (xhigh) xhigh 56%
9 Claude Opus 5 Adaptive Reasoning, Max Effort 56%
10 GPT-5.6 Sol (low) low 55%
11 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 55%
12 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 54%
13 Grok 4.5 high 54%
14 GPT-5.6 Terra max 54%
15 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 54%
16 GPT-5.3 Codex (xhigh) xhigh 53%
17 Gemini 3.5 Flash (high) high 53%
18 Gemini 3.5 Flash (medium) medium 53%
19 Gemini 3.6 Flash (high) high 53%
20 GPT-5.6 Luna max 53%
21 GPT-5.6 Terra (xhigh) xhigh 52%
22 Muse Spark Published configuration 52%
23 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 51%
24 GPT-5.6 Luna (high) high 51%
25 GLM-5.2 (max) max 50%
Show the remaining 225 configurations
#ModelConfigurationPass rate
26 MiMo-V2.5-Pro Published configuration 50%
27 GPT-5.6 Terra (high) high 50%
28 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 50%
29 GPT-5.6 Luna (xhigh) xhigh 50%
30 GPT-5.6 Terra (medium) medium 50%
31 GPT-5.6 Terra (low) low 49%
32 Gemini 3.5 Flash (minimal) minimal 49%
33 Qwen3.7 Max Published configuration 49%
34 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 49%
35 GPT-5.5 Instant (June 2026) June 2026 49%
36 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 48%
37 Hy3 Published configuration 48%
38 Kimi K2.7 Code Published configuration 47%
39 GPT-5.6 Sol (Non-reasoning) Non-reasoning 47%
40 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 46%
41 Inkling (xhigh) xhigh 46%
42 GPT-5.6 Luna (medium) medium 46%
43 GPT-5.6 Luna (low) low 46%
44 Qwen3.7 Plus Published configuration 45%
45 MiniMax-M3 Published configuration 45%
46 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 45%
47 GPT-5.6 Terra (Non-reasoning) Non-reasoning 45%
48 Grok 4.3 (medium) medium 45%
49 Motif 3 (Beta) Beta 44%
50 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 44%
51 Gemma 4 31B (Reasoning) Reasoning 43%
52 Claude 4.5 Haiku (Reasoning) Reasoning 43%
53 MiMo-V2.5 Published configuration 43%
54 Gemini 2.5 Pro Published configuration 43%
55 Nova 2.0 Pro Preview (medium) medium 43%
56 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 42%
57 Ring-2.6-1T Published configuration 42%
58 Agnes 2.5 Pro Alpha Published configuration 42%
59 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 42%
60 Qwen3.5 122B A10B (Reasoning) Reasoning 42%
61 Qwen3.5 397B A17B (Reasoning) Reasoning 42%
62 Gemini 3.1 Flash-Lite Published configuration 42%
63 Grok 4.3 (low) low 42%
64 Nex-N2-Pro Published configuration 42%
65 Hy3-preview (Reasoning) Reasoning 41%
66 Gemma 4 31B (Non-reasoning) Non-reasoning 41%
67 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 41%
68 Cogito v2.1 (Reasoning) Reasoning 41%
69 o3 Published configuration 41%
70 Gemini 3.5 Flash-Lite Published configuration 41%
71 Doubao Seed Code Published configuration 41%
72 Qwen3.6 Plus Published configuration 41%
73 Qwen3.5 Omni Plus Published configuration 41%
74 Gemma 4 26B A4B (Reasoning) Reasoning 40%
75 Step 3.7 Flash Published configuration 40%
76 GPT-5.6 Luna (Non-reasoning) Non-reasoning 40%
77 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 40%
78 Qwen3.6 27B (Reasoning) Reasoning 40%
79 Mistral Medium 3.5 Published configuration 40%
80 Kimi K2.6 (Non-reasoning) Non-reasoning 39%
81 MiMo-V2-Omni-0327 Published configuration 39%
82 Hy3-preview (Non-reasoning) Non-reasoning 39%
83 Magistral Medium 1.2 Published configuration 39%
84 INTELLECT-3 Published configuration 39%
85 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 39%
86 gpt-oss-120b (high) high 39%
87 Qwen3 Next 80B A3B (Reasoning) Reasoning 39%
88 Mercury 2 Published configuration 39%
89 Nova 2.0 Pro Preview (low) low 39%
90 KAT Coder Pro V2 Published configuration 38%
91 MiMo-V2-Flash (Feb 2026) Feb 2026 38%
92 Gemma 4 12B (Reasoning) Reasoning 38%
93 JT-4.1 Flash 236B A21B Published configuration 38%
94 North Mini Code Published configuration 38%
95 Mistral Small 4 (Reasoning) Reasoning 38%
96 Command A+ Published configuration 38%
97 ERNIE 5.0 Thinking Preview Published configuration 38%
98 Grok 4.3 (Non-reasoning) Non-reasoning 37%
99 Apriel-v1.6-15B-Thinker Published configuration 37%
100 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 37%
101 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 37%
102 Qwen3.6 27B (Non-reasoning) Non-reasoning 37%
103 Ling-2.6-1T Published configuration 37%
104 Nova 2.0 Lite (high) high 37%
105 Nova 2.0 Lite (medium) medium 37%
106 MiMo-V2-Omni Published configuration 37%
107 KAT-Coder-Pro V1 Published configuration 37%
108 Mistral Large 3 Published configuration 36%
109 Nova 2.0 Omni (medium) medium 36%
110 GLM-5.2 (Non-reasoning) Non-reasoning 36%
111 Trinity Large Thinking Published configuration 36%
112 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 36%
113 gpt-oss-120b (low) low 36%
114 Qwen3.6 35B A3B (Reasoning) Reasoning 36%
115 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 36%
116 K-EXAONE (Reasoning) Reasoning 36%
117 LongCat 2.0 Published configuration 35%
118 Magistral Small 1.2 Published configuration 35%
119 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 35%
120 Nemotron Cascade 2 30B A3B Published configuration 35%
121 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 35%
122 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 35%
123 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 34%
124 EXAONE 4.0 32B (Reasoning) Reasoning 34%
125 gpt-oss-20b (high) high 34%
126 DiffusionGemma 26B A4B Published configuration 34%
127 Nova 2.0 Omni (low) low 34%
128 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 34%
129 gpt-oss-20b (low) low 34%
130 Nova 2.0 Lite (low) low 33%
131 Mi:dm K 2.5 Pro Published configuration 33%
132 Devstral 2 Published configuration 33%
133 Llama 4 Maverick Published configuration 33%
134 HyperNova 60B 2605 Published configuration 33%
135 K2 Think V2 Published configuration 33%
136 Qwen3 Coder Next Published configuration 32%
137 ERNIE 4.5 300B A47B Published configuration 31%
138 Step3 VL 10B Published configuration 31%
139 Qwen3 Next 80B A3B Instruct Published configuration 31%
140 Qwen3 Omni 30B A3B (Reasoning) Reasoning 31%
141 Llama 3.1 Instruct 405B Published configuration 30%
142 Gemma 4 12B (Non-reasoning) Non-reasoning 30%
143 Mi:dm K 2.5 Pro Preview Published configuration 30%
144 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 30%
145 Olmo 3.1 32B Think Published configuration 29%
146 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 29%
147 JT-35B-Flash Published configuration 29%
148 Devstral Small 2 Published configuration 29%
149 K2-V2 (high) high 29%
150 HyperCLOVA X SEED Think (32B) 32B 28%
151 LongCat Flash Lite Published configuration 28%
152 Motif-2-12.7B-Reasoning Published configuration 28%
153 Command A Published configuration 28%
154 Mistral Small 4 (Non-reasoning) Non-reasoning 28%
155 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 28%
156 EXAONE 4.5 33B Published configuration 28%
157 Nova 2.0 Omni (Non-reasoning) Non-reasoning 28%
158 Nova Premier Published configuration 28%
159 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 28%
160 Qwen3.5 9B (Non-reasoning) Non-reasoning 28%
161 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 28%
162 Qwen3.5 9B (Reasoning) Reasoning 28%
163 JT-MINI Published configuration 27%
164 Ling 2.6 Flash Published configuration 27%
165 K-EXAONE (Non-reasoning) Non-reasoning 27%
166 Solar Open 100B (Reasoning) Reasoning 27%
167 Reka Flash 3 Published configuration 27%
168 Nanbeige4.1-3B Published configuration 27%
169 Sarvam 105B (high) high 26%
170 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 26%
171 Llama 3.3 Instruct 70B Published configuration 26%
172 Phi-4 Published configuration 26%
173 MiMo-V2-Flash (Non-reasoning) Non-reasoning 26%
174 Granite 4.1 30B Published configuration 26%
175 Qwen3.5 Omni Flash Published configuration 25%
176 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 25%
177 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 25%
178 K2-V2 (medium) medium 25%
179 Falcon-H1R-7B Published configuration 25%
180 Solar Pro 3 Published configuration 25%
181 Gemma 4 E4B (Reasoning) Reasoning 24%
182 Llama 3.2 Instruct 90B (Vision) Vision 24%
183 Nova 2.0 Lite (Non-reasoning) Non-reasoning 24%
184 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 24%
185 Ministral 3 14B Published configuration 24%
186 Llama 3.1 Nemotron Instruct 70B Published configuration 23%
187 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 23%
188 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) Non-reasoning 23%
189 K2-V2 (low) low 22%
190 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 22%
191 Granite 4.1 8B Published configuration 22%
192 Olmo 3 7B Think Published configuration 21%
193 Gemma 4 E2B (Reasoning) Reasoning 21%
194 Granite 4.0 H Small Published configuration 21%
195 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 21%
196 Ministral 3 8B Published configuration 21%
197 Gemma 4 E2B (Non-reasoning) Non-reasoning 20%
198 Kimi Linear 48B A3B Instruct Published configuration 20%
199 Sarvam 30B (high) high 19%
200 Jamba 1.7 Large Published configuration 19%
201 Qwen3 Omni 30B A3B Instruct Published configuration 19%
202 Qwen3.5 4B (Non-reasoning) Non-reasoning 18%
203 G9v3-3B Published configuration 18%
204 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 18%
205 Tri-21B-Think Published configuration 17%
206 Llama 4 Scout Published configuration 17%
207 Ring-flash-2.0 Published configuration 17%
208 Olmo 3.1 32B Instruct Published configuration 17%
209 NVIDIA Nemotron 3 Nano 4B Published configuration 16%
210 Qwen3.5 4B (Reasoning) Reasoning 16%
211 Ministral 3 3B Published configuration 14%
212 Ling-mini-2.0 Published configuration 14%
213 Molmo2-8B Published configuration 13%
214 Granite 4.0 Micro Published configuration 12%
215 Granite 4.1 3B Published configuration 12%
216 Llama 3.2 Instruct 11B (Vision) Vision 11%
217 Phi-4 Multimodal Instruct Published configuration 11%
218 LFM2 24B A2B Published configuration 11%
219 Phi-4 Mini Instruct Published configuration 11%
220 Olmo 3 7B Instruct Published configuration 10%
221 Nova Micro Published configuration 9%
222 Exaone 4.0 1.2B (Reasoning) Reasoning 9%
223 Jamba 1.7 Mini Published configuration 9%
224 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) Non-reasoning 9%
225 Granite 4.0 1B Published configuration 9%
226 Granite 4.0 H 1B Published configuration 8%
227 LFM2.5-8B-A1B Published configuration 8%
228 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 7%
229 Qwen3.5 2B (Non-reasoning) Non-reasoning 7%
230 LFM2 8B A1B Published configuration 7%
231 Jamba Reasoning 3B Published configuration 6%
232 Apertus 70B Instruct Published configuration 6%
233 MiniCPM5-1B (Reasoning) Reasoning 4%
234 LFM2.5-1.2B-Thinking Published configuration 4%
235 Apertus 8B Instruct Published configuration 4%
236 Gemma 4 E4B (Non-reasoning) Non-reasoning 4%
237 Molmo 7B-D Published configuration 4%
238 Tiny Aya Global Published configuration 4%
239 LFM2.5-VL-1.6B Published configuration 3%
240 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 3%
241 Qwen3.5 2B (Reasoning) Reasoning 3%
242 LFM2 2.6B Published configuration 3%
243 LFM2.5-1.2B-Instruct Published configuration 2%
244 MiniCPM-V 4.6 1.3B Published configuration 2%
245 Granite 4.0 H 350M Published configuration 2%
246 MiniCPM5-1B (Non-reasoning) Non-reasoning 1%
247 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 1%
248 Granite 4.0 350M Published configuration 1%
249 Gemma 3 270M Published configuration 0%
250 Qwen3.5 0.8B (Reasoning) Reasoning 0%

Showing the top 25 of 250 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A scientific-programming benchmark requiring Python solutions to research-oriented computational problems.

How to read the score

The published metric is Pass rate. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same scicode field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official SciCode resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗