Benchmarks / DeckBench

Measured by Spring Prompt

DeckBench

Given a finished analysis, which models make a deck a manager could present as it is: true to the numbers, telling the right story, and well designed?

Results dated
6 Oct 2026
Models
18
Unit
% of tasks
Licence
Spring Prompt original
Judge
GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.1 Pro, by majority

Draft figure quoted: DeepSeek V4 Pro 0423

Top 15 of 18 results · % of tasks, lower is better. Choose a model to highlight it.Clear highlight

  1. 1 Claude Opus 5.5Anthropic 0.0%
  2. 1 Claude Sonnet 5.5Anthropic 0.0%
  3. 1 Gemini 3.1 Pro PreviewGoogle 0.0%
  4. 1 Kimi K3Moonshot AI 0.0%
  5. 1 GPT-6.1 SolOpenAI 0.0%
  6. 1 GPT-6 AstraOpenAI 0.0%
  7. 1 GPT-6 SolOpenAI 0.0%
  8. 8 Muse Spark 1.3Meta 16.7%
  9. 8 Mistral Large 4Mistral 16.7%
  10. 8 Grok 4.7xAI 16.7%
  11. 11 Qwen3.8 Max (0902)Alibaba 33.3%
  12. 11 Claude Fable 5.1Anthropic 33.3%
  13. 11 DeepSeek V4 Pro 0423DeepSeek 33.3%
  14. 11 GPT-6 LunaOpenAI 33.3%
  15. 15 Claude Haiku 4.5Anthropic 50.0%

Deck rating against cost

What one deck cost at the run date's prices, on a log scale. Up and to the left is better.

0.005001,0001,5002,000 $0.001$0.01$0.1$1$10 Cost per deck, US dollars (log scale) Deck rating GPT-6 Astra GPT-6 Sol GPT-6.1 Sol Claude Sonnet 5.5 Claude Opus 5.5 Claude Fable 5.1

Full results

DeckBench: draft figure quoted, % of tasks, lower is better
#ModelDraft figure quoted
% of tasks, lower is better
Deck rating
rating
Presentable decks
% of tasks
Accurate decks
% of tasks
Design quality
% of the maximum
Cost per deck
US dollars
1 Claude Opus 5.5Anthropic
0.0%
1,35633.3%83.3%72.0%$0.26
1 Claude Sonnet 5.5Anthropic
0.0%
1,35650.0%66.7%78.5%$0.15
1 Gemini 3.1 Pro PreviewGoogle
0.0%
86716.7%50.0%51.4%$0.14
1 Kimi K3Moonshot AI
0.0%
1,2550.0%50.0%81.4%$0.45
1 GPT-6.1 SolOpenAI
0.0%
1,41250.0%100.0%71.4%$0.15
1 GPT-6 AstraOpenAI
0.0%
1,48966.7%100.0%88.5%$1.09
1 GPT-6 SolOpenAI
0.0%
1,46450.0%83.3%87.9%$0.21
8 Muse Spark 1.3Meta
16.7%
8110.0%16.7%56.0%$0.0463
8 Mistral Large 4Mistral
16.7%
5150.0%0.0%30.5%$0.0148
8 Grok 4.7xAI
16.7%
1,24233.3%66.7%76.4%$0.25
11 Qwen3.8 Max (0902)Alibaba
33.3%
1,11233.3%50.0%67.0%$0.15
11 Claude Fable 5.1Anthropic
33.3%
1,30033.3%66.7%66.6%$0.81
11 DeepSeek V4 Pro 0423DeepSeek
33.3%
6150.0%0.0%40.9%$0.0465
11 GPT-6 LunaOpenAI
33.3%
1,2730.0%50.0%78.1%$0.0086
15 Claude Haiku 4.5Anthropic
50.0%
6350.0%16.7%47.6%$0.0461
15 Gemini 3.5 Flash LiteGoogle
50.0%
6440.0%0.0%60.2%$0.0126
15 Gemini 3.8 FlashGoogle
50.0%
9040.0%16.7%61.7%$0.0613
15 Mistral Medium 3.5Mistral
50.0%
6490.0%0.0%43.6%$0.0448

Each model runs at its provider's default reasoning setting, one deck per task. The rating comes from head-to-head comparisons, so it keeps separating models even when several make presentable decks; ranges are 95% intervals from resampling the comparisons.

See the decks

Every model's deck for every public task, as rendered from the slides it returned. Open a deck to see each slide, the judges' rating and what stopped it being presentable.

Hearthside Coffee: Manchester store investment review · for the board of directors (non-executives and the chief executive)
Lumio: Starter churn: what to fund in Q1 2026 · for the leadership team (chief executive, CFO, heads of product, sales and customer success)
Pantry Lane: Paid media: H1 2026 allocation · for the CMO and the finance director
Northgate Homes: Responsive repairs: in-house team or renewed contract · for the board of trustees of a housing association, including two tenant board members
Kestrel Logistics: On-time delivery: what is going wrong and what to fund · for the board and the company's three main investors
Wrenfield Borough Council: Library services: choosing how to save £350,000 · for the council's cabinet (elected members), meeting in public

More from the results

Accurate is not the same as presentable

Share of tasks where the deck said only what the analysis supports, against the share that could also be presented as it is. The gap is layout and design.

AccuratePresentable
GPT-6.1 Sol
100.0% · 50.0%
GPT-6 Astra
100.0% · 66.7%
Claude Opus 5.5
83.3% · 33.3%
GPT-6 Sol
83.3% · 50.0%
Claude Fable 5.1
66.7% · 33.3%
Claude Sonnet 5.5
66.7% · 50.0%
Grok 4.7
66.7% · 33.3%
Qwen3.8 Max (0902)
50.0% · 33.3%
Gemini 3.1 Pro Preview
50.0% · 16.7%
Kimi K3
50.0% · 0.0%
GPT-6 Luna
50.0% · 0.0%
Claude Haiku 4.5
16.7% · 0.0%
Gemini 3.8 Flash
16.7% · 0.0%
Muse Spark 1.3
16.7% · 0.0%
DeepSeek V4 Pro 0423
0.0% · 0.0%
Gemini 3.5 Flash Lite
0.0% · 0.0%
Mistral Large 4
0.0% · 0.0%
Mistral Medium 3.5
0.0% · 0.0%

Why decks are not presentable

Share of tasks failing each check. A deck can fail several at once; any one means it needs fixing before it is shown.

ModelLayout defectsSlides needing workDraft figure quotedUnsupported numbersUnsupported claimsCaveat droppedMisleading metricFindings missingRecommendation
GPT-6 AstraOpenAI0.0%33.3%0.0%0.0%0.0%0.0%0.0%0.0%0.0%
GPT-6 SolOpenAI33.3%0.0%0.0%0.0%0.0%0.0%0.0%16.7%0.0%
GPT-6.1 SolOpenAI16.7%50.0%0.0%0.0%0.0%0.0%0.0%0.0%0.0%
Claude Sonnet 5.5Anthropic0.0%33.3%0.0%0.0%0.0%16.7%0.0%0.0%16.7%
Claude Opus 5.5Anthropic16.7%66.7%0.0%0.0%0.0%16.7%0.0%0.0%0.0%
Claude Fable 5.1Anthropic66.7%50.0%33.3%0.0%0.0%0.0%0.0%0.0%0.0%
GPT-6 LunaOpenAI33.3%16.7%33.3%0.0%0.0%0.0%0.0%0.0%16.7%
Kimi K3Moonshot AI66.7%16.7%0.0%0.0%33.3%0.0%0.0%16.7%33.3%
Grok 4.7xAI16.7%16.7%16.7%0.0%0.0%16.7%0.0%0.0%16.7%
Qwen3.8 Max (0902)Alibaba33.3%50.0%33.3%0.0%16.7%16.7%0.0%0.0%0.0%
Gemini 3.8 FlashGoogle83.3%33.3%50.0%16.7%50.0%0.0%0.0%0.0%33.3%
Gemini 3.1 Pro PreviewGoogle50.0%83.3%0.0%0.0%0.0%16.7%0.0%0.0%50.0%
Muse Spark 1.3Meta100.0%50.0%16.7%0.0%16.7%16.7%0.0%0.0%66.7%
Mistral Medium 3.5Mistral33.3%50.0%50.0%0.0%33.3%0.0%0.0%0.0%50.0%
Gemini 3.5 Flash LiteGoogle66.7%16.7%50.0%0.0%16.7%16.7%0.0%33.3%83.3%
Claude Haiku 4.5Anthropic50.0%66.7%50.0%0.0%33.3%0.0%0.0%16.7%50.0%
DeepSeek V4 Pro 0423DeepSeek66.7%83.3%33.3%0.0%66.7%16.7%0.0%16.7%66.7%
Mistral Large 4Mistral66.7%100.0%16.7%33.3%50.0%16.7%0.0%0.0%83.3%

How DeckBench works

  1. 1

    The analysis

    An invented company's memo with findings, a recommendation and caveats, two to four data tables, the audience, the decision and a brand kit.

  2. 2

    The deck

    The model returns every slide as positioned text, shapes, native charts and tables. We render it to PowerPoint and to images.

  3. 3

    The checks

    Code checks the layout and traces every number to the analysis; three judges check the story and rate each slide's design.

  4. 4

    Head to head

    On each task the judges compare decks in pairs and pick the one they would rather present. The votes become the rating.

The task

Turning an analysis into slides is one of the most common things people ask AI to do at work, and the risk is not ugly slides: it is a number that is not in the analysis, a caveat that falls off, or a draft figure that ends up in front of the board.

Each task is an invented company's analysis written for the benchmark, with planted traps: a caveat that must travel with its number, a figure the memo says not to quote, a metric it warns is misleading, and a finding that contradicts a senior stakeholder's view.

What makes a deck presentable

Layout
Nothing off the slide or colliding, body text at least 12 points, and text readable against its background. Text wraps from real font metrics, so overflow is measured.
Numbers
Every number on a slide, in a chart or in a table must come from the analysis; numbers code cannot match go to the judges.
Story
The key findings, the recommendation within the first two slides, caveats with their numbers, and no misleading metric used as evidence.
Design
Each slide rated polished, acceptable or needs work. A presentable deck has no slide that needs work.

Why a rating as well

Presentable is a bar: once the best models clear it on every task, it stops telling them apart. The head-to-head rating does not run out of room, because a deck only has to be preferred to another one. Each deck meets about six others per task; a 400-point gap means the higher-rated model's deck is preferred about ten times out of eleven.

What it measures

  • Every number traced to the analysis
  • The story: findings, recommendation and caveats
  • Layout and legibility, checked by code
  • Design quality, judged slide by slide

What it does not measure

  • Your templates, data or brand guidelines
  • Decks built in PowerPoint or Google Slides by an agent clicking through the app
  • Speaking notes or delivery

Method

  • Layout, legibility and numbers are checked by code
  • Three judges from three vendors decide each judged check by majority
  • No judge compares decks from its own vendor
  • Two tasks are held back so the set can be refreshed

Checking the judge

A single judge was the strictest of three on unsupported claims (it flagged 24 material claims in a pilot where the other two flagged 12 and 1), and its own model led the pilot. So every judged check is decided by a majority of three judges: GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.1 Pro. In the pilot they agreed with each other on 97 to 100% of the checks on findings, numbers, recommendations and misleading metrics, and on 89 to 90% of decisions about whether a slide needs work. In head-to-head comparisons a judge never votes on a pair that includes its own vendor's deck. A sample of human ratings is still to come.

Failures

Failures count against a model: a product with no usable output is a failed listing. They are listed here so you can see why.

Not ranked

  • GLM 5.3: no usable deck: reply cut off at the token limit on 6 of 6 tasks