Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

CritPt benchmark: model scores and methodology

A research-level physics benchmark using composite reasoning challenges and executable or symbolic answer formats.

Current published leader
GPT-5.6 Sol · max
Top score
32%
Primary metric
Accuracy · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current CritPt model scores

246 configurations · 178 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 246 current configurations with a reported CritPt score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on CritPt

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. GPT-5.6 Sol 32%
  2. GPT-5.5 Pro xhigh 31%
  3. GPT-5.6 Terra 30%
  4. Claude Opus 5 29%
  5. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 29%
  6. Gemini 3 Deep Think 26%
  7. Kimi K3 23%
  8. GLM-5.2 max 21%
  9. GPT-5.6 Luna 21%
  10. Gemini 3.1 Pro (Preview) 18%
  11. Claude Sonnet 5 Adaptive Reasoning, Max Effort 17%
  12. GPT-5.3 Codex xhigh 17%
  13. Grok 4.5 15%
  14. Muse Spark 1.1 xhigh 15%
  15. Qwen3.7 Max 13%
  16. Gemini 3.5 Flash high 13%
  17. DeepSeek V4 Pro Reasoning, Max Effort 13%
  18. Muse Spark 11%
  19. Agnes 2.5 Pro Alpha 11%
  20. Gemini 3.6 Flash high 11%
0.0 16.1 32.3

Accuracy · higher is better

Current configurations with a reported CritPt value; missing values are omitted.

How to interpret the result

What do CritPt results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher accuracy is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read CritPt as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationAccuracy
1 GPT-5.6 SolLeader max 32%
2 GPT-5.5 Pro (xhigh) xhigh 31%
3 GPT-5.6 Terra max 30%
4 Claude Opus 5 Adaptive Reasoning, Max Effort 29%
5 GPT-5.6 Sol (xhigh) xhigh 29%
6 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 29%
7 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 28%
8 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 28%
9 GPT-5.6 Terra (xhigh) xhigh 27%
10 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 27%
11 GPT-5.6 Sol (high) high 26%
12 Gemini 3 Deep Think Published configuration 26%
13 Kimi K3 Published configuration 23%
14 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 23%
15 GPT-5.6 Sol (medium) medium 23%
16 GPT-5.6 Terra (high) high 23%
17 GLM-5.2 (max) max 21%
18 GPT-5.6 Luna max 21%
19 GPT-5.6 Luna (xhigh) xhigh 21%
20 Gemini 3.1 Pro (Preview) Published configuration 18%
21 GPT-5.6 Terra (medium) medium 17%
22 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 17%
23 GPT-5.3 Codex (xhigh) xhigh 17%
24 GPT-5.6 Luna (high) high 17%
25 Grok 4.5 high 15%
Show the remaining 221 configurations
#ModelConfigurationAccuracy
26 Muse Spark 1.1 (xhigh) xhigh 15%
27 GPT-5.6 Sol (low) low 15%
28 Qwen3.7 Max Published configuration 13%
29 Gemini 3.5 Flash (high) high 13%
30 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 13%
31 Muse Spark Published configuration 11%
32 Agnes 2.5 Pro Alpha Published configuration 11%
33 Gemini 3.5 Flash (medium) medium 11%
34 Gemini 3.6 Flash (high) high 11%
35 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 10%
36 Kimi K2.7 Code Published configuration 10%
37 GPT-5.6 Terra (low) low 9%
38 Qwen3.7 Plus Published configuration 9%
39 Nex-N2-Pro Published configuration 9%
40 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 7%
41 Motif 3 (Beta) Beta 6%
42 Inkling (xhigh) xhigh 5%
43 GPT-5.6 Sol (Non-reasoning) Non-reasoning 5%
44 GPT-5.6 Luna (medium) medium 5%
45 Grok 4.3 (medium) medium 5%
46 Hy3 Published configuration 5%
47 Hy3-preview (Reasoning) Reasoning 5%
48 MiMo-V2.5-Pro Published configuration 4%
49 MiMo-V2.5 Published configuration 4%
50 MiniMax-M3 Published configuration 4%
51 Ring-2.6-1T Published configuration 4%
52 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 3%
53 GLM-5.2 (Non-reasoning) Non-reasoning 3%
54 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 3%
55 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 3%
56 MiMo-V2-Flash (Feb 2026) Feb 2026 3%
57 Qwen3.6 Plus Published configuration 3%
58 GPT-5.6 Luna (low) low 3%
59 Gemini 2.5 Pro Published configuration 3%
60 LongCat 2.0 Published configuration 3%
61 Step 3.7 Flash Published configuration 2%
62 GPT-5.6 Terra (Non-reasoning) Non-reasoning 2%
63 Qwen3.5 397B A17B (Reasoning) Reasoning 2%
64 ERNIE 5.0 Thinking Preview Published configuration 1%
65 Gemini 3.5 Flash (minimal) minimal 1%
66 Gemma 4 31B (Reasoning) Reasoning 1%
67 Kimi K2.6 (Non-reasoning) Non-reasoning 1%
68 gpt-oss-20b (high) high 1%
69 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 1%
70 Gemini 3.1 Flash-Lite Published configuration 1%
71 MiMo-V2-Omni Published configuration 1%
72 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 1%
73 Qwen3.6 27B (Reasoning) Reasoning 1%
74 gpt-oss-120b (high) high 1%
75 o3 Published configuration 1%
76 K-EXAONE (Reasoning) Reasoning 1%
77 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 1%
78 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 1%
79 HyperNova 60B 2605 Published configuration 1%
80 MiMo-V2-Omni-0327 Published configuration 1%
81 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 1%
82 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 1%
83 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 1%
84 Qwen3.6 27B (Non-reasoning) Non-reasoning 1%
85 Trinity Large Thinking Published configuration 1%
86 Mercury 2 Published configuration 1%
87 Gemma 4 E4B (Reasoning) Reasoning 1%
88 Grok 4.3 (low) low 1%
89 Nemotron Cascade 2 30B A3B Published configuration 1%
90 Qwen3.5 122B A10B (Reasoning) Reasoning 1%
91 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 1%
92 Qwen3.5 9B (Non-reasoning) Non-reasoning 1%
93 Qwen3.5 Omni Plus Published configuration 1%
94 Apriel-v1.6-15B-Thinker Published configuration 0%
95 Command A+ Published configuration 0%
96 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 0%
97 DiffusionGemma 26B A4B Published configuration 0%
98 Doubao Seed Code Published configuration 0%
99 EXAONE 4.5 33B Published configuration 0%
100 Falcon-H1R-7B Published configuration 0%
101 GPT-5.6 Luna (Non-reasoning) Non-reasoning 0%
102 Gemma 4 E4B (Non-reasoning) Non-reasoning 0%
103 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 0%
104 Hy3-preview (Non-reasoning) Non-reasoning 0%
105 INTELLECT-3 Published configuration 0%
106 Ling-2.6-1T Published configuration 0%
107 Magistral Medium 1.2 Published configuration 0%
108 Magistral Small 1.2 Published configuration 0%
109 Mistral Small 4 (Non-reasoning) Non-reasoning 0%
110 Mistral Small 4 (Reasoning) Reasoning 0%
111 North Mini Code Published configuration 0%
112 Nova 2.0 Lite (high) high 0%
113 Qwen3.5 9B (Reasoning) Reasoning 0%
114 Qwen3.6 35B A3B (Reasoning) Reasoning 0%
115 Sarvam 30B (high) high 0%
116 Tri-21B-Think Published configuration 0%
117 Ring-flash-2.0 Published configuration 0%
118 Apertus 70B Instruct Published configuration 0%
119 Apertus 8B Instruct Published configuration 0%
120 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 0%
121 Claude 4.5 Haiku (Reasoning) Reasoning 0%
122 Cogito v2.1 (Reasoning) Reasoning 0%
123 Command A Published configuration 0%
124 Devstral 2 Published configuration 0%
125 Devstral Small 2 Published configuration 0%
126 ERNIE 4.5 300B A47B Published configuration 0%
127 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 0%
128 EXAONE 4.0 32B (Reasoning) Reasoning 0%
129 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 0%
130 Exaone 4.0 1.2B (Reasoning) Reasoning 0%
131 G9v3-3B Published configuration 0%
132 GPT-5.5 Instant (June 2026) June 2026 0%
133 Gemini 3.5 Flash-Lite Published configuration 0%
134 Gemma 3 270M Published configuration 0%
135 Gemma 4 12B (Non-reasoning) Non-reasoning 0%
136 Gemma 4 12B (Reasoning) Reasoning 0%
137 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 0%
138 Gemma 4 26B A4B (Reasoning) Reasoning 0%
139 Gemma 4 31B (Non-reasoning) Non-reasoning 0%
140 Gemma 4 E2B (Non-reasoning) Non-reasoning 0%
141 Gemma 4 E2B (Reasoning) Reasoning 0%
142 Granite 4.0 1B Published configuration 0%
143 Granite 4.0 350M Published configuration 0%
144 Granite 4.0 H 1B Published configuration 0%
145 Granite 4.0 H 350M Published configuration 0%
146 Granite 4.0 H Small Published configuration 0%
147 Granite 4.0 Micro Published configuration 0%
148 Granite 4.1 30B Published configuration 0%
149 Granite 4.1 3B Published configuration 0%
150 Granite 4.1 8B Published configuration 0%
151 Grok 4.3 (Non-reasoning) Non-reasoning 0%
152 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 0%
153 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 0%
154 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 0%
155 HyperCLOVA X SEED Think (32B) 32B 0%
156 JT-35B-Flash Published configuration 0%
157 JT-4.1 Flash 236B A21B Published configuration 0%
158 JT-MINI Published configuration 0%
159 Jamba 1.7 Large Published configuration 0%
160 Jamba 1.7 Mini Published configuration 0%
161 Jamba Reasoning 3B Published configuration 0%
162 K-EXAONE (Non-reasoning) Non-reasoning 0%
163 K2 Think V2 Published configuration 0%
164 K2-V2 (high) high 0%
165 K2-V2 (low) low 0%
166 K2-V2 (medium) medium 0%
167 KAT Coder Pro V2 Published configuration 0%
168 KAT-Coder-Pro V1 Published configuration 0%
169 LFM2 2.6B Published configuration 0%
170 LFM2 24B A2B Published configuration 0%
171 LFM2 8B A1B Published configuration 0%
172 LFM2.5-1.2B-Instruct Published configuration 0%
173 LFM2.5-1.2B-Thinking Published configuration 0%
174 LFM2.5-8B-A1B Published configuration 0%
175 LFM2.5-VL-1.6B Published configuration 0%
176 Ling 2.6 Flash Published configuration 0%
177 Ling-mini-2.0 Published configuration 0%
178 Llama 3.1 Instruct 405B Published configuration 0%
179 Llama 3.1 Nemotron Instruct 70B Published configuration 0%
180 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 0%
181 Llama 3.2 Instruct 11B (Vision) Vision 0%
182 Llama 3.3 Instruct 70B Published configuration 0%
183 Llama 4 Maverick Published configuration 0%
184 Llama 4 Scout Published configuration 0%
185 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 0%
186 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 0%
187 LongCat Flash Lite Published configuration 0%
188 Mi:dm K 2.5 Pro Published configuration 0%
189 Mi:dm K 2.5 Pro Preview Published configuration 0%
190 MiMo-V2-Flash (Non-reasoning) Non-reasoning 0%
191 MiniCPM-V 4.6 1.3B Published configuration 0%
192 MiniCPM5-1B (Non-reasoning) Non-reasoning 0%
193 MiniCPM5-1B (Reasoning) Reasoning 0%
194 Ministral 3 14B Published configuration 0%
195 Ministral 3 3B Published configuration 0%
196 Ministral 3 8B Published configuration 0%
197 Mistral Large 3 Published configuration 0%
198 Mistral Medium 3.5 Published configuration 0%
199 Molmo2-8B Published configuration 0%
200 Motif-2-12.7B-Reasoning Published configuration 0%
201 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 0%
202 NVIDIA Nemotron 3 Nano 4B Published configuration 0%
203 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 0%
204 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 0%
205 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 0%
206 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 0%
207 Nanbeige4.1-3B Published configuration 0%
208 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 0%
209 Nova 2.0 Lite (Non-reasoning) Non-reasoning 0%
210 Nova 2.0 Lite (low) low 0%
211 Nova 2.0 Lite (medium) medium 0%
212 Nova 2.0 Omni (Non-reasoning) Non-reasoning 0%
213 Nova 2.0 Omni (low) low 0%
214 Nova 2.0 Omni (medium) medium 0%
215 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 0%
216 Nova 2.0 Pro Preview (low) low 0%
217 Nova 2.0 Pro Preview (medium) medium 0%
218 Nova Micro Published configuration 0%
219 Nova Premier Published configuration 0%
220 Olmo 3 7B Instruct Published configuration 0%
221 Olmo 3 7B Think Published configuration 0%
222 Olmo 3.1 32B Instruct Published configuration 0%
223 Olmo 3.1 32B Think Published configuration 0%
224 Phi-4 Published configuration 0%
225 Phi-4 Mini Instruct Published configuration 0%
226 Qwen3 Coder Next Published configuration 0%
227 Qwen3 Next 80B A3B (Reasoning) Reasoning 0%
228 Qwen3 Next 80B A3B Instruct Published configuration 0%
229 Qwen3 Omni 30B A3B (Reasoning) Reasoning 0%
230 Qwen3 Omni 30B A3B Instruct Published configuration 0%
231 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 0%
232 Qwen3.5 0.8B (Reasoning) Reasoning 0%
233 Qwen3.5 2B (Non-reasoning) Non-reasoning 0%
234 Qwen3.5 2B (Reasoning) Reasoning 0%
235 Qwen3.5 4B (Non-reasoning) Non-reasoning 0%
236 Qwen3.5 4B (Reasoning) Reasoning 0%
237 Qwen3.5 Omni Flash Published configuration 0%
238 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 0%
239 Reka Flash 3 Published configuration 0%
240 Sarvam 105B (high) high 0%
241 Solar Open 100B (Reasoning) Reasoning 0%
242 Solar Pro 3 Published configuration 0%
243 Step3 VL 10B Published configuration 0%
244 Tiny Aya Global Published configuration 0%
245 gpt-oss-120b (low) low 0%
246 gpt-oss-20b (low) low 0%

Showing the top 25 of 246 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A research-level physics benchmark using composite reasoning challenges and executable or symbolic answer formats.

How to read the score

The published metric is Accuracy. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same critpt field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official CritPt resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗