Ranks follow the score as shown, so equal numbers share a rank. Each model runs at its provider's default reasoning setting, one deck per task. The rating comes from head-to-head comparisons, so it keeps separating models even when several make presentable decks; ranges are 95% intervals from resampling the comparisons.
See the decks
Every model's deck for every public task, as rendered from the slides it returned. Open a deck to see each slide, the judges' rating and what stopped it being presentable.
Hearthside Coffee: Manchester store investment review · for the board of directors (non-executives and the chief executive)
Northgate Homes: Responsive repairs: in-house team or renewed contract · for the board of trustees of a housing association, including two tenant board members
Share of tasks where the deck said only what the analysis supports, against the share that could also be presented as it is. The gap is layout and design.
AccuratePresentable
GPT-6.1 Sol
100.0% · 50.0%
GPT-6 Astra
100.0% · 66.7%
Claude Opus 5.5
83.3% · 33.3%
GPT-6 Sol
83.3% · 50.0%
Claude Fable 5.1
66.7% · 33.3%
Claude Sonnet 5.5
66.7% · 50.0%
Grok 4.7
66.7% · 33.3%
Qwen3.8-Max (0902)
50.0% · 33.3%
Gemini 3.1 Pro Preview
50.0% · 16.7%
Kimi K3
50.0% · 0.0%
GPT-6 Luna
50.0% · 0.0%
Claude Haiku 5.5
33.3% · 33.3%
Claude Haiku 4.5
16.7% · 0.0%
Gemini 3.8 Flash
16.7% · 0.0%
Muse Spark 1.3
16.7% · 0.0%
DeepSeek-V4-Pro (0423)
0.0% · 0.0%
Gemini 3.5 Flash-Lite
0.0% · 0.0%
Mistral Large 4
0.0% · 0.0%
Mistral Medium 3.5
0.0% · 0.0%
0%25%50%75%100%
Why decks are not presentable
Share of tasks failing each check. A deck can fail several at once; any one means it needs fixing before it is shown.
Model
Layout defects
Slides needing work
Draft figure quoted
Unsupported numbers
Unsupported claims
Caveat dropped
Misleading metric
Findings missing
Recommendation
GPT-6 AstraOpenAI
0.0%
33.3%
0.0%
0.0%
0.0%
0.0%
0.0%
0.0%
0.0%
GPT-6 SolOpenAI
33.3%
0.0%
0.0%
0.0%
0.0%
0.0%
0.0%
16.7%
0.0%
GPT-6.1 SolOpenAI
16.7%
50.0%
0.0%
0.0%
0.0%
0.0%
0.0%
0.0%
0.0%
Claude Sonnet 5.5Anthropic
0.0%
33.3%
0.0%
0.0%
0.0%
16.7%
0.0%
0.0%
16.7%
Claude Opus 5.5Anthropic
16.7%
66.7%
0.0%
0.0%
0.0%
16.7%
0.0%
0.0%
0.0%
Claude Fable 5.1Anthropic
66.7%
50.0%
33.3%
0.0%
0.0%
0.0%
0.0%
0.0%
0.0%
GPT-6 LunaOpenAI
33.3%
16.7%
33.3%
0.0%
0.0%
0.0%
0.0%
0.0%
16.7%
Kimi K3Moonshot AI
66.7%
16.7%
0.0%
0.0%
33.3%
0.0%
0.0%
16.7%
33.3%
Grok 4.7xAI
16.7%
16.7%
16.7%
0.0%
0.0%
16.7%
0.0%
0.0%
16.7%
Claude Haiku 5.5Anthropic
33.3%
16.7%
33.3%
16.7%
33.3%
16.7%
16.7%
16.7%
33.3%
Qwen3.8-Max (0902)Alibaba
33.3%
50.0%
33.3%
0.0%
16.7%
16.7%
0.0%
0.0%
0.0%
Gemini 3.8 FlashGoogle
83.3%
33.3%
50.0%
16.7%
50.0%
0.0%
0.0%
0.0%
33.3%
Gemini 3.1 Pro PreviewGoogle
50.0%
83.3%
0.0%
0.0%
0.0%
16.7%
0.0%
0.0%
50.0%
Muse Spark 1.3Meta
100.0%
50.0%
16.7%
0.0%
16.7%
16.7%
0.0%
0.0%
66.7%
Claude Haiku 4.5Anthropic
50.0%
66.7%
50.0%
0.0%
33.3%
0.0%
0.0%
16.7%
50.0%
Mistral Medium 3.5Mistral AI
33.3%
50.0%
50.0%
0.0%
33.3%
0.0%
0.0%
0.0%
50.0%
Gemini 3.5 Flash-LiteGoogle
66.7%
16.7%
50.0%
0.0%
16.7%
16.7%
0.0%
33.3%
83.3%
DeepSeek-V4-Pro (0423)DeepSeek
66.7%
83.3%
33.3%
0.0%
66.7%
16.7%
0.0%
16.7%
66.7%
Mistral Large 4Mistral AI
66.7%
100.0%
16.7%
33.3%
50.0%
16.7%
0.0%
0.0%
83.3%
Every task, every model
Whether each deck could be presented as it is, and if not, how many kinds of issue it had. Each square opens the deck.
✓ presentable as it is · a number: the kinds of issue that stopped it · – no usable deck.
How DeckBench works
1
The analysis
An invented company's memo with findings, a recommendation and caveats, two to four data tables, the audience, the decision and a brand kit.
2
The deck
The model returns every slide as positioned text, shapes, native charts and tables. We render it to PowerPoint and to images.
3
The checks
Code checks the layout and traces every number to the analysis; three judges check the story and rate each slide's design.
4
Head to head
On each task the judges compare decks in pairs and pick the one they would rather present. The votes become the rating.
The task
Turning an analysis into slides is one of the most common things people ask AI to do at work, and the risk is not ugly slides: it is a number that is not in the analysis, a caveat that falls off, or a draft figure that ends up in front of the board.
Each task is an invented company's analysis written for the benchmark, with planted traps: a caveat that must travel with its number, a figure the memo says not to quote, a metric it warns is misleading, and a finding that contradicts a senior stakeholder's view.
What makes a deck presentable
Layout
Nothing off the slide or colliding, body text at least 12 points, and text readable against its background. Text wraps from real font metrics, so overflow is measured.
Numbers
Every number on a slide, in a chart or in a table must come from the analysis; numbers code cannot match go to the judges.
Story
The key findings, the recommendation within the first two slides, caveats with their numbers, and no misleading metric used as evidence.
Design
Each slide rated polished, acceptable or needs work. A presentable deck has no slide that needs work.
Why a rating as well
Presentable is a bar: once the best models clear it on every task, it stops telling them apart. The head-to-head rating does not run out of room, because a deck only has to be preferred to another one. Each deck meets about six others per task; a 400-point gap means the higher-rated model's deck is preferred about ten times out of eleven.
What it measures
Every number traced to the analysis
The story: findings, recommendation and caveats
Layout and legibility, checked by code
Design quality, judged slide by slide
What it does not measure
Your templates, data or brand guidelines
Decks built in PowerPoint or Google Slides by an agent clicking through the app
Speaking notes or delivery
Method
Layout, legibility and numbers are checked by code
Three judges from three vendors decide each judged check by majority
No judge compares decks from its own vendor
Two tasks are held back so the set can be refreshed
Checking the judge
A single judge was the strictest of three on unsupported claims (it flagged 24 material claims in a pilot where the other two flagged 12 and 1), and its own model led the pilot. So every judged check is decided by a majority of three judges: GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.1 Pro. In the pilot they agreed with each other on 97 to 100% of the checks on findings, numbers, recommendations and misleading metrics, and on 89 to 90% of decisions about whether a slide needs work. In head-to-head comparisons a judge never votes on a pair that includes its own vendor's deck. A sample of human ratings is still to come.
Failures
A model that made no usable deck on any task has no rating, so it is listed here instead of ranked. They are listed here so you can see why.
Not ranked
GLM 5.3: no usable deck: reply cut off at the token limit on 6 of 6 tasks