Confirm Action

Are you sure you want to proceed?

Artificial Analysis · Public leaderboard snapshot

Results through 2026-07-30

MMMU-Pro benchmark: model scores and methodology

A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.

Current published leader
Claude Opus 5 · Adaptive Reasoning, Max Effort
Top score
85%
Primary metric
Accuracy · higher is better
Evaluation subject
Model configuration in Artificial Analysis's evaluation system

Artificial Analysis public leaderboard snapshot

All current MMMU-Pro model scores

121 configurations · 75 model entries

This reviewed Artificial Analysis snapshot · 30 July 2026 contains 121 current configurations with a reported MMMU-Pro score. Figures are the rounded values displayed by Artificial Analysis; underlying source precision is retained in SQLite.

Leading models on MMMU-Pro

The chart shows the 20 highest current configurations in this one Artificial Analysis snapshot. One mark per model, at its strongest published configuration. These results sit in a narrow band, so the scale below is zoomed — read each mark against the labelled axis, not the left edge. Marks are placed on the full source precision, so two models sharing a rounded label can still sit at different points. Score labels reproduce Artificial Analysis's rounded display value.

  1. Claude Opus 5 85%
  2. Gemini 3.5 Flash high 84%
  3. GPT-5.6 Sol 83%
  4. Gemini 3.6 Flash high 83%
  5. Gemini 3.1 Pro (Preview) 82%
  6. GPT-5.6 Terra 81%
  7. Kimi K3 81%
  8. Muse Spark 81%
  9. Qwen3.7 Plus 80%
  10. Grok 4.5 80%
  11. Gemini 3.5 Flash-Lite 79%
  12. GPT-5.6 Luna 79%
  13. MiniMax-M3 79%
  14. GPT-5.3 Codex xhigh 78%
  15. Qwen3.6 Plus 78%
  16. Claude Sonnet 5 Adaptive Reasoning, Max Effort 77%
  17. Qwen3.5 397B A17B Reasoning 77%
  18. Grok 4.3 medium 76%
  19. Gemini 3.1 Flash-Lite 76%
  20. MiMo-V2.5 75%
72.2 79.0 85.9

Accuracy · higher is better · zoomed scale

Current configurations with a reported MMMU-Pro value; missing values are omitted.

How to interpret the result

What do MMMU-Pro results mean?

Treat each row as a result for the named model configuration inside Artificial Analysis's evaluation setup, not as a property of bare model weights.

1. Read the displayed figure

Higher accuracy is better. The public table rounds the displayed score, while Springprompt retains the source precision.

2. Check the evaluated system

The score depends on the model configuration, Artificial Analysis harness, tools, task budget, repeats, and scoring protocol.

3. Compare within one contract

Use rows from this same field and snapshot for the cleanest comparison. Do not merge vendor-reported or differently harnessed scores into this table.

The leaderboard is decision evidence, not a universal model ranking: match the benchmark contract to the work you actually need done.

Read MMMU-Pro as a result of Artificial Analysis's evaluated model configuration and methodology—not as a context-free model property.

Every row comes from Artificial Analysis snapshot · 30 July 2026, recorded 30 Jul 2026.

#ModelConfigurationAccuracy
1 Claude Opus 5Leader Adaptive Reasoning, Max Effort 85%
2 Gemini 3.5 Flash (high) high 84%
3 Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Adaptive Reasoning, Xhigh Effort 84%
4 Gemini 3.5 Flash (medium) medium 84%
5 GPT-5.6 Sol max 83%
6 Gemini 3.6 Flash (high) high 83%
7 GPT-5.6 Sol (xhigh) xhigh 83%
8 Claude Opus 5 (Adaptive Reasoning, High Effort) Adaptive Reasoning, High Effort 82%
9 Gemini 3.1 Pro (Preview) Published configuration 82%
10 GPT-5.6 Sol (high) high 82%
11 Claude Opus 5 (Adaptive Reasoning, Medium Effort) Adaptive Reasoning, Medium Effort 82%
12 GPT-5.6 Sol (medium) medium 81%
13 GPT-5.6 Sol (low) low 81%
14 GPT-5.6 Terra max 81%
15 Kimi K3 Published configuration 81%
16 Muse Spark Published configuration 81%
17 Qwen3.7 Plus Published configuration 80%
18 Grok 4.5 high 80%
19 Gemini 3.5 Flash (minimal) minimal 80%
20 Claude Opus 5 (Adaptive Reasoning, Low Effort) Adaptive Reasoning, Low Effort 80%
21 GPT-5.6 Terra (xhigh) xhigh 79%
22 GPT-5.6 Terra (high) high 79%
23 Gemini 3.5 Flash-Lite Published configuration 79%
24 GPT-5.6 Luna max 79%
25 GPT-5.6 Luna (xhigh) xhigh 79%
Show the remaining 96 configurations
#ModelConfigurationAccuracy
26 MiniMax-M3 Published configuration 79%
27 GPT-5.3 Codex (xhigh) xhigh 78%
28 Qwen3.6 Plus Published configuration 78%
29 GPT-5.6 Luna (high) high 78%
30 Claude Sonnet 5 (Adaptive Reasoning, Max Effort) Adaptive Reasoning, Max Effort 77%
31 Qwen3.5 397B A17B (Reasoning) Reasoning 77%
32 GPT-5.6 Terra (medium) medium 77%
33 GPT-5.6 Terra (low) low 76%
34 GPT-5.6 Luna (medium) medium 76%
35 Grok 4.3 (medium) medium 76%
36 Gemini 3.1 Flash-Lite Published configuration 76%
37 MiMo-V2.5 Published configuration 75%
38 Step 3.7 Flash Published configuration 75%
39 Qwen3.6 35B A3B (Reasoning) Reasoning 75%
40 Qwen3.5 122B A10B (Reasoning) Reasoning 75%
41 Gemini 2.5 Pro Published configuration 75%
42 Qwen3.6 27B (Reasoning) Reasoning 75%
43 GPT-5.6 Luna (low) low 74%
44 MiMo-V2-Omni-0327 Published configuration 74%
45 Inkling (xhigh) xhigh 73%
46 Gemma 4 31B (Reasoning) Reasoning 73%
47 Grok 4.3 (low) low 73%
48 Claude Sonnet 5 (Non-reasoning, High Effort) Non-reasoning, High Effort 72%
49 GPT-5.6 Sol (Non-reasoning) Non-reasoning 72%
50 Qwen3.6 27B (Non-reasoning) Non-reasoning 72%
51 Qwen3.6 35B A3B (Non-reasoning) Non-reasoning 71%
52 Qwen3.5 Omni Plus Published configuration 71%
53 Gemma 4 31B (Non-reasoning) Non-reasoning 70%
54 Qwen3.5 122B A10B (Non-reasoning) Non-reasoning 70%
55 o3 Published configuration 70%
56 MiMo-V2-Omni Published configuration 70%
57 Gemma 4 12B (Reasoning) Reasoning 70%
58 Gemma 4 26B A4B (Reasoning) Reasoning 69%
59 Qwen3.5 9B (Reasoning) Reasoning 69%
60 Claude Sonnet 4.6 (Non-reasoning, Low Effort) Non-reasoning, Low Effort 69%
61 Qwen3.5 35B A3B (Non-reasoning) Non-reasoning 69%
62 Doubao Seed Code Published configuration 68%
63 EXAONE 4.5 33B Published configuration 67%
64 Qwen3.5 9B (Non-reasoning) Non-reasoning 67%
65 GPT-5.6 Terra (Non-reasoning) Non-reasoning 67%
66 Gemma 4 26B A4B (Non-reasoning) Non-reasoning 67%
67 DiffusionGemma 26B A4B Published configuration 67%
68 Qwen3.5 4B (Reasoning) Reasoning 65%
69 Mistral Medium 3.5 Published configuration 65%
70 Grok 4.3 (Non-reasoning) Non-reasoning 65%
71 Qwen3.5 Omni Flash Published configuration 65%
72 ERNIE 5.0 Thinking Preview Published configuration 65%
73 Nova 2.0 Pro Preview (medium) medium 65%
74 JT-4.1 Flash 236B A21B Published configuration 64%
75 Step3 VL 10B Published configuration 64%
76 Nova 2.0 Lite (high) high 64%
77 Command A+ Published configuration 63%
78 Nova 2.0 Pro Preview (low) low 63%
79 Nova 2.0 Lite (medium) medium 63%
80 Llama 4 Maverick Published configuration 62%
81 Qwen3.5 4B (Non-reasoning) Non-reasoning 62%
82 Gemma 4 12B (Non-reasoning) Non-reasoning 62%
83 Nova 2.0 Omni (medium) medium 62%
84 GPT-5.6 Luna (Non-reasoning) Non-reasoning 60%
85 Qwen3 Omni 30B A3B (Reasoning) Reasoning 60%
86 Nova 2.0 Omni (low) low 60%
87 Magistral Medium 1.2 Published configuration 60%
88 Claude 4.5 Haiku (Reasoning) Reasoning 59%
89 Nova 2.0 Lite (low) low 58%
90 Mistral Small 4 (Reasoning) Reasoning 57%
91 Mistral Large 3 Published configuration 56%
92 Magistral Small 1.2 Published configuration 55%
93 Qwen3 Omni 30B A3B Instruct Published configuration 55%
94 Claude 4.5 Haiku (Non-reasoning) Non-reasoning 55%
95 Nemotron 3 Nano Omni 30B A3B Reasoning Published configuration 53%
96 Llama 4 Scout Published configuration 53%
97 NVIDIA Nemotron Nano 12B v2 VL (Reasoning) Reasoning 53%
98 Qwen3.5 397B A17B (Non-reasoning) Non-reasoning 53%
99 Gemma 4 E4B (Reasoning) Reasoning 51%
100 Gemma 4 E4B (Non-reasoning) Non-reasoning 51%
101 Nova 2.0 Omni (Non-reasoning) Non-reasoning 50%
102 Ministral 3 14B Published configuration 50%
103 Nova 2.0 Lite (Non-reasoning) Non-reasoning 49%
104 Mistral Small 4 (Non-reasoning) Non-reasoning 46%
105 Ministral 3 8B Published configuration 46%
106 Devstral Small 2 Published configuration 45%
107 Gemma 4 E2B (Reasoning) Reasoning 45%
108 NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning) Non-reasoning 45%
109 Qwen3.5 2B (Reasoning) Reasoning 43%
110 Qwen3.5 2B (Non-reasoning) Non-reasoning 43%
111 Gemma 4 E2B (Non-reasoning) Non-reasoning 42%
112 Llama 3.2 Instruct 90B (Vision) Vision 39%
113 Ministral 3 3B Published configuration 38%
114 MiniCPM-V 4.6 1.3B Published configuration 38%
115 Molmo2-8B Published configuration 37%
116 Llama 3.2 Instruct 11B (Vision) Vision 29%
117 LFM2.5-VL-1.6B Published configuration 27%
118 Qwen3.5 0.8B (Reasoning) Reasoning 26%
119 Qwen3.5 0.8B (Non-reasoning) Non-reasoning 26%
120 Molmo 7B-D Published configuration 25%
121 Phi-4 Multimodal Instruct Published configuration 15%

Showing the top 25 of 121 published configurations.

What Artificial Analysis held constantPublished harness, scoring, and budget boundaries Open contract
Comparison source
Artificial Analysis public LLM leaderboard captured Artificial Analysis snapshot · 30 July 2026.
Harness
Artificial Analysis's independently operated benchmark implementation for this metric.
What varies
The published model configuration and provider-side implementation; reasoning variants remain separate rows.
Tools
Tool access follows the metric-specific Artificial Analysis methodology and is not assumed to be uniform across different benchmarks.
Budget
Task counts, repeats, turn limits, and timeouts follow the cited methodology; they are not equal-compute guarantees across model providers.
Comparison limit
Comparable within this source field and snapshot; not interchangeable with scores from another harness or protocol version.

What this benchmark tests

A multimodal academic reasoning benchmark designed to reduce shortcuts and guessing across many disciplines.

How to read the score

The published metric is Accuracy. Springprompt reproduces Artificial Analysis's rounded public-table figure and preserves its underlying numeric value for provenance.

A missing source value is not scored as zero: that configuration is omitted from this benchmark page.

Comparability policy

Why this is a system evaluation

A row identifies the model configuration, but the measured subject also includes the evaluator's prompts, harness, tools, budgets, repeats, and grader.

That is why Springprompt does not combine these figures with vendor claims or results from another implementation simply because the benchmark name looks similar.

What can be compared here

Every row on this page comes from the same Artificial Analysis snapshot · 30 July 2026 leaderboard payload and the same mmmuPro field.

Estimated Intelligence Index rows remain visible but carry an explicit estimate label. Missing fields and deprecated models are not manufactured into pages or zero scores.

Official MMMU-Pro resources 2 links · show

Go deeper

Turn benchmark evidence into a model decision

Browse Spring Prompt’s task-level model evidence, compare the published configurations above, or join the product waitlist to build an evaluation around your own workflow.

Sources and provenance

Spring Prompt stores a reviewed, content-addressed evidence manifest for every citation. The linked official source remains canonical.

  1. 1.Artificial Analysis public LLM leaderboard ↗Artificial Analysis · model scores and source display values · retrieved 2026-07-30 · evidence 4e8ccd3759d9
  2. 2.Artificial Analysis intelligence benchmarking methodology ↗Artificial Analysis · methodology and evaluation-contract interpretation · retrieved 2026-07-30 · evidence 45ccc8609f26

Read the official scoring methodology ↗