Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

Humanity's Last Exam benchmark: model scores and methodology

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

Current published leader
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Top score
53%3-way tie
Primary metric
Accuracy · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current Humanity's Last Exam model scores

250 configurations · 182 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 250 current configurations with a reported Humanity's Last Exam score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on Humanity's Last Exam

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. Score labels reproduce Artificial Analysis's rounded display value.

  1. Claude Fable 5 Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 53%
  2. Claude Opus 5 53%
  3. GPT-5.6 Sol 47%
  4. Muse Spark 1.1 xhigh 45%
  5. Gemini 3.1 Pro (Preview) 45%
  6. Kimi K3 44%
  7. GPT-5.6 Terra 42%
  8. Gemini 3.5 Flash high 41%
  9. Grok 4.5 40%
  10. GLM-5.2 max 40%
  11. Muse Spark 40%
  12. GPT-5.3 Codex xhigh 40%
  13. Claude Sonnet 5 Adaptive Reasoning, Max Effort 40%
  14. Gemini 3.6 Flash high 38%
  15. Motif 3 Beta 38%
  16. Qwen3.7 Max 38%
  17. GPT-5.6 Luna 37%
  18. MiniMax-M3 37%
  19. DeepSeek V4 Pro Reasoning, Max Effort 36%
  20. MiMo-V2.5-Pro 34%
0.0 26.7 53.3

Accuracy · higher is better

Current configurations with a reported Humanity's Last Exam value; missing values are omitted.

How to interpret the result

What do Humanity's Last Exam results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher accuracy is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read Humanity's Last Exam as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationAccuracy
1 Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)Leader Adaptive Reasoning, Max Effort, Opus 4.8 Fallback 53%
2 Claude Opus 5 Adaptive Reasoning, Max Effort 53%
3 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 53%
4 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 51%
5 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 49%
6 GPT-5.6 Sol max 47%
7 Muse Spark 1.1 (xhigh) xhigh 45%
8 GPT-5.6 Sol (xhigh) xhigh 45%
9 Gemini 3.1 Pro (Preview) Published configuration 45%
10 Kimi K3 Published configuration 44%
11 GPT-5.6 Sol (high) high 44%
12 GPT-5.6 Terra max 42%
13 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 41%
14 Gemini 3.5 Flash (high) high 41%
15 Grok 4.5 high 40%
16 GLM-5.2 (max) max 40%
17 GPT-5.6 Terra (xhigh) xhigh 40%
18 Muse Spark Published configuration 40%
19 GPT-5.3 Codex (xhigh) xhigh 40%
20 Gemini 3.5 Flash (medium) medium 40%
21 GPT-5.6 Sol (medium) medium 40%
22 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 40%
23 Gemini 3.6 Flash (high) high 38%
24 Motif 3 (Beta) Beta 38%
25 Qwen3.7 Max Published configuration 38%
Show the remaining 225 configurations
#ModelConfigurationAccuracy
26 GPT-5.6 Luna max 37%
27 MiniMax-M3 Published configuration 37%
28 GPT-5.6 Terra (high) high 37%
29 GPT-5.6 Sol (low) low 37%
30 DeepSeek V4 Pro (Reasoning, Max Effort) Reasoning, Max Effort 36%
31 GPT-5.6 Luna (xhigh) xhigh 36%
32 MiMo-V2.5-Pro Published configuration 34%
33 DeepSeek V4 Pro (Reasoning, High Effort) Reasoning, High Effort 34%
34 KAT-Coder-Pro V1 Published configuration 33%
35 Qwen3.7 Plus Published configuration 33%
36 Kimi K2.7 Code Published configuration 33%
37 Nex-N2-Pro Published configuration 32%
38 DeepSeek V4 Flash (Reasoning, Max Effort) Reasoning, Max Effort 32%
39 LongCat 2.0 Published configuration 32%
40 Agnes 2.5 Pro Alpha Published configuration 32%
41 Hy3 Published configuration 32%
42 GPT-5.6 Luna (high) high 32%
43 GPT-5.6 Terra (medium) medium 32%
44 Inkling (xhigh) xhigh 30%
45 Grok 4.3 (medium) medium 28%
46 DeepSeek V4 Flash (Reasoning, High Effort) Reasoning, High Effort 28%
47 GPT-5.6 Terra (low) low 27%
48 Qwen3.5 397B A17B (Reasoning) Reasoning 27%
49 Nemotron 3 Ultra 550B A55B (Reasoning) Reasoning 27%
50 Qwen3.6 Plus Published configuration 26%
51 Hy3-preview (Reasoning) Reasoning 26%
52 MiMo-V2.5 Published configuration 25%
53 GPT-5.6 Luna (medium) medium 24%
54 Qwen3.5 122B A10B (Reasoning) Reasoning 23%
55 Gemini 3.5 Flash (minimal) minimal 23%
56 Gemma 4 31B (Reasoning) Reasoning 23%
57 Qwen3.6 27B (Reasoning) Reasoning 22%
58 Gemini 2.5 Pro Published configuration 21%
59 MiMo-V2-Omni-0327 Published configuration 20%
60 Qwen3.6 35B A3B (Reasoning) Reasoning 20%
61 o3 Published configuration 20%
62 MiMo-V2-Flash (Feb 2026) Feb 2026 20%
63 MiMo-V2-Omni Published configuration 20%
64 Step 3.7 Flash Published configuration 20%
65 NVIDIA Nemotron 3 Super 120B A12B (Reasoning) Reasoning 19%
66 GPT-5.6 Luna (low) low 19%
67 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 19%
68 GPT-5.5 Instant (June 2026) June 2026 19%
69 gpt-oss-120b (high) high 18%
70 Gemma 4 26B A4B (Reasoning) Reasoning 18%
71 Ring-2.6-1T Published configuration 18%
72 Kimi K2.6 (Non-reasoning) Non-reasoning 18%
73 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 18%
74 Gemini 3.5 Flash-Lite Published configuration 18%
75 Grok 4.3 (low) low 17%
76 Gemini 3.1 Flash-Lite Published configuration 16%
77 JT-4.1 Flash 236B A21B Published configuration 16%
78 KAT Coder Pro V2 Published configuration 16%
79 GPT-5.6 Sol (Non-reasoning) Non-reasoning 16%
80 Mercury 2 Published configuration 16%
81 HyperNova 60B 2605 Published configuration 15%
82 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 15%
83 Gemma 4 12B (Reasoning) Reasoning 15%
84 Trinity Large Thinking Published configuration 15%
85 Qwen3.5 Omni Plus Published configuration 14%
86 Qwen3.6 27B (Non-reasoning) Non-reasoning 14%
87 MiMo-V2.5-Pro (Non-reasoning) Non-reasoning 13%
88 Qwen3.5 9B (Reasoning) Reasoning 13%
89 Doubao Seed Code Published configuration 13%
90 K-EXAONE (Reasoning) Reasoning 13%
91 Mistral Medium 3.5 Published configuration 13%
92 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 13%
93 ERNIE 5.0 Thinking Preview Published configuration 13%
94 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 13%
95 INTELLECT-3 Published configuration 12%
96 Qwen3 Next 80B A3B (Reasoning) Reasoning 12%
97 EXAONE 4.5 33B Published configuration 12%
98 Gemma 4 31B (Non-reasoning) Non-reasoning 11%
99 Command A+ Published configuration 11%
100 Nemotron Cascade 2 30B A3B Published configuration 11%
101 GPT-5.6 Terra (Non-reasoning) Non-reasoning 11%
102 Cogito v2.1 (Reasoning) Reasoning 11%
103 Nova 2.0 Lite (high) high 11%
104 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 11%
105 Falcon-H1R-7B Published configuration 11%
106 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 11%
107 EXAONE 4.0 32B (Reasoning) Reasoning 11%
108 Hermes 4 - Llama-3.1 405B (Reasoning) Reasoning 10%
109 NVIDIA Nemotron 3 Nano 30B A3B (Reasoning) Reasoning 10%
110 DiffusionGemma 26B A4B Published configuration 10%
111 Step3 VL 10B Published configuration 10%
112 Sarvam 105B (high) high 10%
113 Solar Pro 3 Published configuration 10%
114 Nanbeige4.1-3B Published configuration 10%
115 North Mini Code Published configuration 10%
116 K2-V2 (high) high 10%
117 Apriel-v1.6-15B-Thinker Published configuration 10%
118 gpt-oss-20b (high) high 10%
119 Claude 4.5 Haiku (Reasoning) Reasoning 10%
120 Magistral Medium 1.2 Published configuration 10%
121 Mistral Small 4 (Reasoning) Reasoning 9%
122 K2 Think V2 Published configuration 9%
123 Qwen3 Coder Next Published configuration 9%
124 Solar Open 100B (Reasoning) Reasoning 9%
125 Nova 2.0 Pro Preview (medium) medium 9%
126 Ring-flash-2.0 Published configuration 9%
127 Mi:dm K 2.5 Pro Preview Published configuration 9%
128 Nova 2.0 Lite (medium) medium 9%
129 Qwen3.5 9B (Non-reasoning) Non-reasoning 9%
130 GLM-5.2 (Non-reasoning) Non-reasoning 8%
131 Ling-2.6-1T Published configuration 8%
132 Motif-2-12.7B-Reasoning Published configuration 8%
133 Llama 3.1 Nemotron Ultra 253B v1 (Reasoning) Reasoning 8%
134 MiMo-V2-Flash (Non-reasoning) Non-reasoning 8%
135 Hermes 4 - Llama-3.1 70B (Reasoning) Reasoning 8%
136 Qwen3.5 4B (Reasoning) Reasoning 8%
137 DeepSeek V4 Pro (Non-reasoning) Non-reasoning 8%
138 Mi:dm K 2.5 Pro Published configuration 8%
139 Qwen3.5 4B (Non-reasoning) Non-reasoning 7%
140 Qwen3 Next 80B A3B Instruct Published configuration 7%
141 Qwen3 Omni 30B A3B (Reasoning) Reasoning 7%
142 Qwen3.5 Omni Flash Published configuration 7%
143 Sarvam 30B (high) high 7%
144 DeepSeek V4 Flash (Non-reasoning) Non-reasoning 7%
145 LFM2.5-8B-A1B Published configuration 7%
146 Llama Nemotron Super 49B v1.5 (Reasoning) Reasoning 7%
147 LFM2.5-1.2B-Instruct Published configuration 7%
148 Nova 2.0 Omni (medium) medium 7%
149 GPT-5.6 Luna (Non-reasoning) Non-reasoning 7%
150 JT-MINI Published configuration 7%
151 MiniCPM5-1B (Reasoning) Reasoning 7%
152 Grok 4.3 (Non-reasoning) Non-reasoning 6%
153 Granite 4.0 H 350M Published configuration 6%
154 Hy3-preview (Non-reasoning) Non-reasoning 6%
155 Gemma 4 12B (Non-reasoning) Non-reasoning 6%
156 Ling 2.6 Flash Published configuration 6%
157 JT-35B-Flash Published configuration 6%
158 Tri-21B-Think Published configuration 6%
159 LFM2.5-1.2B-Thinking Published configuration 6%
160 Magistral Small 1.2 Published configuration 6%
161 LongCat Flash Lite Published configuration 6%
162 Olmo 3.1 32B Think Published configuration 6%
163 Exaone 4.0 1.2B (Non-reasoning) Non-reasoning 6%
164 Exaone 4.0 1.2B (Reasoning) Reasoning 6%
165 Olmo 3 7B Instruct Published configuration 6%
166 Granite 4.0 350M Published configuration 6%
167 Olmo 3 7B Think Published configuration 6%
168 Apertus 70B Instruct Published configuration 5%
169 HyperCLOVA X SEED Think (32B) 32B 5%
170 K-EXAONE (Non-reasoning) Non-reasoning 5%
171 Ministral 3 3B Published configuration 5%
172 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 5%
173 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 5%
174 LFM2 2.6B Published configuration 5%
175 Tiny Aya Global Published configuration 5%
176 gpt-oss-120b (low) low 5%
177 Nova 2.0 Pro Preview (low) low 5%
178 Llama 3.2 Instruct 11B (Vision) Vision 5%
179 Granite 4.0 1B Published configuration 5%
180 LFM2.5-VL-1.6B Published configuration 5%
181 Granite 4.0 Micro Published configuration 5%
182 Qwen3 Omni 30B A3B Instruct Published configuration 5%
183 Reka Flash 3 Published configuration 5%
184 Molmo 7B-D Published configuration 5%
185 gpt-oss-20b (low) low 5%
186 Granite 4.0 H 1B Published configuration 5%
187 Ling-mini-2.0 Published configuration 5%
188 Apertus 8B Instruct Published configuration 5%
189 Llama 3.2 Instruct 90B (Vision) Vision 5%
190 MiniCPM-V 4.6 1.3B Published configuration 5%
191 Olmo 3.1 32B Instruct Published configuration 5%
192 EXAONE 4.0 32B (Non-reasoning) Non-reasoning 5%
193 LFM2 8B A1B Published configuration 5%
194 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 5%
195 Qwen3.5 2B (Non-reasoning) Non-reasoning 5%
196 Llama 4 Maverick Published configuration 5%
197 Gemma 4 E2B (Reasoning) Reasoning 5%
198 NVIDIA Nemotron 3 Nano 4B Published configuration 5%
199 Gemma 4 E4B (Non-reasoning) Non-reasoning 5%
200 Nova Premier Published configuration 5%
201 Nova Micro Published configuration 5%
202 Ministral 3 14B Published configuration 5%
203 NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) Non-reasoning 5%
204 Llama 3.1 Nemotron Instruct 70B Published configuration 5%
205 Jamba Reasoning 3B Published configuration 5%
206 MiniCPM5-1B (Non-reasoning) Non-reasoning 5%
207 NVIDIA Nemotron Nano 9B V2 (Reasoning) Reasoning 5%
208 Command A Published configuration 5%
209 Gemma 4 E2B (Non-reasoning) Non-reasoning 4%
210 Jamba 1.7 Mini Published configuration 4%
211 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 4%
212 LFM2 24B A2B Published configuration 4%
213 Molmo2-8B Published configuration 4%
214 K2-V2 (medium) medium 4%
215 Phi-4 Multimodal Instruct Published configuration 4%
216 Llama 4 Scout Published configuration 4%
217 Llama Nemotron Super 49B v1.5 (Non-reasoning) Non-reasoning 4%
218 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 4%
219 Ministral 3 8B Published configuration 4%
220 DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning) Non-reasoning 4%
221 Llama 3.1 Instruct 405B Published configuration 4%
222 Granite 4.1 30B Published configuration 4%
223 Nova 2.0 Lite (low) low 4%
224 Phi-4 Mini Instruct Published configuration 4%
225 Gemma 3 270M Published configuration 4%
226 Hermes 4 - Llama-3.1 405B (Non-reasoning) Non-reasoning 4%
227 Mistral Large 3 Published configuration 4%
228 Phi-4 Published configuration 4%
229 G9v3-3B Published configuration 4%
230 NVIDIA Nemotron Nano 9B V2 (Non-reasoning) Non-reasoning 4%
231 Nova 2.0 Omni (low) low 4%
232 Llama 3.3 Instruct 70B Published configuration 4%
233 Nova 2.0 Pro Preview (Non-reasoning) Non-reasoning 4%
234 DeepHermes 3 - Mistral 24B Preview (Non-reasoning) Non-reasoning 4%
235 K2-V2 (low) low 4%
236 Nova 2.0 Omni (Non-reasoning) Non-reasoning 4%
237 Granite 4.1 8B Published configuration 4%
238 Jamba 1.7 Large Published configuration 4%
239 Gemma 4 E4B (Reasoning) Reasoning 4%
240 Granite 4.0 H Small Published configuration 4%
241 Mistral Small 4 (Non-reasoning) Non-reasoning 4%
242 Hermes 4 - Llama-3.1 70B (Non-reasoning) Non-reasoning 4%
243 Devstral 2 Published configuration 4%
244 ERNIE 4.5 300B A47B Published configuration 4%
245 Devstral Small 2 Published configuration 3%
246 Granite 4.1 3B Published configuration 3%
247 Nova 2.0 Lite (Non-reasoning) Non-reasoning 3%
248 Kimi Linear 48B A3B Instruct Published configuration 3%
249 Qwen3.5 2B (Reasoning) Reasoning 2%
250 Qwen3.5 0.8B (Reasoning) Reasoning 1%

Showing the top 25 of 250 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A broad expert-level benchmark of difficult academic reasoning and knowledge questions.

How to read the score

The published metric is Accuracy. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same hle field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official Humanity's Last Exam resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗