Benchmarks / DeckBench

Measured by Spring Prompt

DeckBench

Given a finished analysis, which models make a deck a manager could present as it is: true to the numbers, telling the right story, and well designed?

Last updated 8 Oct 2026

Results dated
6 Oct 2026 to 8 Oct 2026
Models
19
Unit
US dollars
Licence
Spring Prompt original
Judge
GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.1 Pro, by majority · bias check
Runs
1 deck per task, 6 tasks

Cost per deck: Claude Haiku 5.5

Top 15 of 19 results · US dollars, lower is better. Choose a model to highlight it.Clear highlight

  1. 1 GPT-6 LunaOpenAI $0.0087
  2. 2 Claude Haiku 5.5Anthropic $0.0091
  3. 3 Gemini 3.5 Flash-LiteGoogle $0.0126
  4. 4 Mistral Large 4Mistral AI $0.0148
  5. 5 Mistral Medium 3.5Mistral AI $0.0448
  6. 6 Claude Haiku 4.5Anthropic $0.0461
  7. 7 Muse Spark 1.3Meta $0.0463
  8. 8 DeepSeek-V4-Pro (0423)DeepSeek $0.0465
  9. 9 Gemini 3.8 FlashGoogle $0.0613
  10. 10 Gemini 3.1 Pro PreviewGoogle $0.14
  11. 11 Qwen3.8-Max (0902)Alibaba $0.15
  12. 11 Claude Sonnet 5.5Anthropic $0.15
  13. 11 GPT-6.1 SolOpenAI $0.15
  14. 14 GPT-6 SolOpenAI $0.21
  15. 15 Grok 4.7xAI $0.25

Deck rating against cost

What one deck cost at the run date's prices, on a log scale. Up and to the left is better.

Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost

05001,0001,5002,000 $0.001$0.01$0.1$1$10 Cost per deck, US dollars (log scale) Deck rating GPT-6 Astra GPT-6 Sol GPT-6.1 Sol Claude Sonnet 5.5 Claude Opus 5.5 Claude Fable 5.1 GPT-6 Luna

Best per budget

  1. GPT-6 Luna: 1,259 at $0.0087
  2. Claude Sonnet 5.5: 1,352 at $0.15
  3. GPT-6.1 Sol: 1,398 at $0.15
  4. GPT-6 Sol: 1,450 at $0.21
  5. GPT-6 Astra: 1,477 at $1.09

Cheapest first: each model here beats every cheaper one on score.

Full results

DeckBench: cost per deck, US dollars, lower is better
#ModelCost per deck
US dollars, lower is better
Deck rating
rating
Presentable decks
% of tasks
Accurate decks
% of tasks
Design quality
% of the maximum
1 GPT-6 LunaOpenAI
$0.0087
1,2590.0%
0 of 6
50.0%
3 of 6
78.1%
2 Claude Haiku 5.5Anthropic
$0.0091
1,13033.3%
2 of 6
33.3%
2 of 6
61.1%
3 Gemini 3.5 Flash-LiteGoogle
$0.0126
6450.0%
0 of 6
0.0%
0 of 6
60.2%
4 Mistral Large 4Mistral AI
$0.0148
5160.0%
0 of 6
0.0%
0 of 6
30.5%
5 Mistral Medium 3.5Mistral AI
$0.0448
6460.0%
0 of 6
0.0%
0 of 6
43.7%
6 Claude Haiku 4.5Anthropic
$0.0461
6600.0%
0 of 6
16.7%
1 of 6
47.6%
7 Muse Spark 1.3Meta
$0.0463
8210.0%
0 of 6
16.7%
1 of 6
56.1%
8 DeepSeek-V4-Pro (0423)DeepSeek
$0.0465
6130.0%
0 of 6
0.0%
0 of 6
40.9%
9 Gemini 3.8 FlashGoogle
$0.0613
8940.0%
0 of 6
16.7%
1 of 6
61.7%
10 Gemini 3.1 Pro PreviewGoogle
$0.14
85816.7%
1 of 6
50.0%
3 of 6
51.4%
11 Qwen3.8-Max (0902)Alibaba
$0.15
1,10033.3%
2 of 6
50.0%
3 of 6
67.0%
11 Claude Sonnet 5.5Anthropic
$0.15
1,35250.0%
3 of 6
66.7%
4 of 6
78.5%
11 GPT-6.1 SolOpenAI
$0.15
1,39850.0%
3 of 6
100.0%
6 of 6
71.4%
14 GPT-6 SolOpenAI
$0.21
1,45050.0%
3 of 6
83.3%
5 of 6
87.9%
15 Grok 4.7xAI
$0.25
1,22933.3%
2 of 6
66.7%
4 of 6
76.4%
16 Claude Opus 5.5Anthropic
$0.26
1,35033.3%
2 of 6
83.3%
5 of 6
72.0%
17 Kimi K3Moonshot AI
$0.45
1,2290.0%
0 of 6
50.0%
3 of 6
81.4%
18 Claude Fable 5.1Anthropic
$0.81
1,27233.3%
2 of 6
66.7%
4 of 6
66.6%
19 GPT-6 AstraOpenAI
$1.09
1,47766.7%
4 of 6
100.0%
6 of 6
88.5%

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model runs at its provider's default reasoning setting, one deck per task. The rating comes from head-to-head comparisons, so it keeps separating models even when several make presentable decks; ranges are 95% intervals from resampling the comparisons.

See the decks

Every model's deck for every public task, as rendered from the slides it returned. Open a deck to see each slide, the judges' rating and what stopped it being presentable.

Hearthside Coffee: Manchester store investment review · for the board of directors (non-executives and the chief executive)
Compare decks side by side
Lumio: Starter churn: what to fund in Q1 2026 · for the leadership team (chief executive, CFO, heads of product, sales and customer success)
Compare decks side by side
Pantry Lane: Paid media: H1 2026 allocation · for the CMO and the finance director
Compare decks side by side
Northgate Homes: Responsive repairs: in-house team or renewed contract · for the board of trustees of a housing association, including two tenant board members
Compare decks side by side
Kestrel Logistics: On-time delivery: what is going wrong and what to fund · for the board and the company's three main investors
Compare decks side by side
Wrenfield Borough Council: Library services: choosing how to save £350,000 · for the council's cabinet (elected members), meeting in public
Compare decks side by side

More from the results

Accurate is not the same as presentable

Share of tasks where the deck said only what the analysis supports, against the share that could also be presented as it is. The gap is layout and design.

AccuratePresentable
GPT-6.1 Sol
100.0% · 50.0%
GPT-6 Astra
100.0% · 66.7%
Claude Opus 5.5
83.3% · 33.3%
GPT-6 Sol
83.3% · 50.0%
Claude Fable 5.1
66.7% · 33.3%
Claude Sonnet 5.5
66.7% · 50.0%
Grok 4.7
66.7% · 33.3%
Qwen3.8-Max (0902)
50.0% · 33.3%
Gemini 3.1 Pro Preview
50.0% · 16.7%
Kimi K3
50.0% · 0.0%
GPT-6 Luna
50.0% · 0.0%
Claude Haiku 5.5
33.3% · 33.3%
Claude Haiku 4.5
16.7% · 0.0%
Gemini 3.8 Flash
16.7% · 0.0%
Muse Spark 1.3
16.7% · 0.0%
DeepSeek-V4-Pro (0423)
0.0% · 0.0%
Gemini 3.5 Flash-Lite
0.0% · 0.0%
Mistral Large 4
0.0% · 0.0%
Mistral Medium 3.5
0.0% · 0.0%

Why decks are not presentable

Share of tasks failing each check. A deck can fail several at once; any one means it needs fixing before it is shown.

ModelLayout defectsSlides needing workDraft figure quotedUnsupported numbersUnsupported claimsCaveat droppedMisleading metricFindings missingRecommendation
GPT-6 AstraOpenAI0.0%33.3%0.0%0.0%0.0%0.0%0.0%0.0%0.0%
GPT-6 SolOpenAI33.3%0.0%0.0%0.0%0.0%0.0%0.0%16.7%0.0%
GPT-6.1 SolOpenAI16.7%50.0%0.0%0.0%0.0%0.0%0.0%0.0%0.0%
Claude Sonnet 5.5Anthropic0.0%33.3%0.0%0.0%0.0%16.7%0.0%0.0%16.7%
Claude Opus 5.5Anthropic16.7%66.7%0.0%0.0%0.0%16.7%0.0%0.0%0.0%
Claude Fable 5.1Anthropic66.7%50.0%33.3%0.0%0.0%0.0%0.0%0.0%0.0%
GPT-6 LunaOpenAI33.3%16.7%33.3%0.0%0.0%0.0%0.0%0.0%16.7%
Kimi K3Moonshot AI66.7%16.7%0.0%0.0%33.3%0.0%0.0%16.7%33.3%
Grok 4.7xAI16.7%16.7%16.7%0.0%0.0%16.7%0.0%0.0%16.7%
Claude Haiku 5.5Anthropic33.3%16.7%33.3%16.7%33.3%16.7%16.7%16.7%33.3%
Qwen3.8-Max (0902)Alibaba33.3%50.0%33.3%0.0%16.7%16.7%0.0%0.0%0.0%
Gemini 3.8 FlashGoogle83.3%33.3%50.0%16.7%50.0%0.0%0.0%0.0%33.3%
Gemini 3.1 Pro PreviewGoogle50.0%83.3%0.0%0.0%0.0%16.7%0.0%0.0%50.0%
Muse Spark 1.3Meta100.0%50.0%16.7%0.0%16.7%16.7%0.0%0.0%66.7%
Claude Haiku 4.5Anthropic50.0%66.7%50.0%0.0%33.3%0.0%0.0%16.7%50.0%
Mistral Medium 3.5Mistral AI33.3%50.0%50.0%0.0%33.3%0.0%0.0%0.0%50.0%
Gemini 3.5 Flash-LiteGoogle66.7%16.7%50.0%0.0%16.7%16.7%0.0%33.3%83.3%
DeepSeek-V4-Pro (0423)DeepSeek66.7%83.3%33.3%0.0%66.7%16.7%0.0%16.7%66.7%
Mistral Large 4Mistral AI66.7%100.0%16.7%33.3%50.0%16.7%0.0%0.0%83.3%

Every task, every model

Whether each deck could be presented as it is, and if not, how many kinds of issue it had. Each square opens the deck.

ModelHearthside CoffeeLumioPantry LaneNorthgate HomesKestrel LogisticsWrenfield Borough CouncilPresentable
GPT-6 Astra✓✓✓1✓14 of 6
GPT-6 Sol1✓1✓✓13 of 6
GPT-6.1 Sol✓211✓✓3 of 6
Claude Sonnet 5.5✓1✓✓123 of 6
Claude Opus 5.5✓11✓222 of 6
Claude Fable 5.13✓321✓2 of 6
Grok 4.7✓1✓1122 of 6
Claude Haiku 5.52–✓1✓12 of 6
Qwen3.8 Max (0902)23✓1✓32 of 6
Gemini 3.1 Pro Preview21✓3331 of 6
GPT-6 Luna1111110 of 6
Kimi K35111110 of 6
Gemini 3.8 Flash2251420 of 6
Muse Spark 1.32214340 of 6
Claude Haiku 4.52323240 of 6
Mistral Medium 3.51114330 of 6
Gemini 3.5 Flash Lite4342130 of 6
DeepSeek V4 Pro 04232444430 of 6
Mistral Large 43525430 of 6
GLM 5.3––––––0 of 6

✓ presentable as it is · a number: the kinds of issue that stopped it · – no usable deck.

How DeckBench works

  1. 1

    The analysis

    An invented company's memo with findings, a recommendation and caveats, two to four data tables, the audience, the decision and a brand kit.

  2. 2

    The deck

    The model returns every slide as positioned text, shapes, native charts and tables. We render it to PowerPoint and to images.

  3. 3

    The checks

    Code checks the layout and traces every number to the analysis; three judges check the story and rate each slide's design.

  4. 4

    Head to head

    On each task the judges compare decks in pairs and pick the one they would rather present. The votes become the rating.

The task

Turning an analysis into slides is one of the most common things people ask AI to do at work, and the risk is not ugly slides: it is a number that is not in the analysis, a caveat that falls off, or a draft figure that ends up in front of the board.

Each task is an invented company's analysis written for the benchmark, with planted traps: a caveat that must travel with its number, a figure the memo says not to quote, a metric it warns is misleading, and a finding that contradicts a senior stakeholder's view.

What makes a deck presentable

Layout
Nothing off the slide or colliding, body text at least 12 points, and text readable against its background. Text wraps from real font metrics, so overflow is measured.
Numbers
Every number on a slide, in a chart or in a table must come from the analysis; numbers code cannot match go to the judges.
Story
The key findings, the recommendation within the first two slides, caveats with their numbers, and no misleading metric used as evidence.
Design
Each slide rated polished, acceptable or needs work. A presentable deck has no slide that needs work.

Why a rating as well

Presentable is a bar: once the best models clear it on every task, it stops telling them apart. The head-to-head rating does not run out of room, because a deck only has to be preferred to another one. Each deck meets about six others per task; a 400-point gap means the higher-rated model's deck is preferred about ten times out of eleven.

What it measures

  • Every number traced to the analysis
  • The story: findings, recommendation and caveats
  • Layout and legibility, checked by code
  • Design quality, judged slide by slide

What it does not measure

  • Your templates, data or brand guidelines
  • Decks built in PowerPoint or Google Slides by an agent clicking through the app
  • Speaking notes or delivery

Method

  • Layout, legibility and numbers are checked by code
  • Three judges from three vendors decide each judged check by majority
  • No judge compares decks from its own vendor
  • Two tasks are held back so the set can be refreshed

Checking the judge

A single judge was the strictest of three on unsupported claims (it flagged 24 material claims in a pilot where the other two flagged 12 and 1), and its own model led the pilot. So every judged check is decided by a majority of three judges: GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.1 Pro. In the pilot they agreed with each other on 97 to 100% of the checks on findings, numbers, recommendations and misleading metrics, and on 89 to 90% of decisions about whether a slide needs work. In head-to-head comparisons a judge never votes on a pair that includes its own vendor's deck. A sample of human ratings is still to come.

Failures

A model that made no usable deck on any task has no rating, so it is listed here instead of ranked. They are listed here so you can see why.

Not ranked

  • GLM 5.3: no usable deck: reply cut off at the token limit on 6 of 6 tasks