Benchmarks / ROASBench

Measured by Spring Prompt

ROASBench

Given twelve months of paid-marketing decisions for a simulated brand, which models grow the business and which overspend?

Last updated 8 Oct 2026

Results dated
29 Sep 2026 to 8 Oct 2026
Results
20 configurations of 19 models
Unit
score out of 100
Licence
Spring Prompt original
Runs
1 per model

Score: GPT-6 Astra

Top 15 of 19 results · score out of 100, higher is better · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ GPT-6 AstraOpenAI 56.6
  2. 2≈ GPT-6 SolOpenAI 54.1
  3. 3≈ GPT-6 LunaOpenAI 53.4
  4. 4≈ Claude Sonnet 5.5Anthropic 52.0
  5. 5≈ Claude Opus 5.5Anthropic 51.7
  6. 6≈ Claude Fable 5.1Anthropic 51.5
  7. 7≈ Gemini 3.8 FlashGoogle 51.4
  8. 8 Kimi K3Moonshot AI 44.1
  9. 9 Gemini 3.1 Pro PreviewGoogle 43.5
  10. 10 Mistral Large 4Mistral AI 37.0
  11. 11 Grok 4.7xAI 36.6
  12. 12 Muse Spark 1.3Meta 34.1
  13. 13 DeepSeek-V4-Pro (0423)DeepSeek 28.3
  14. 14 Qwen3.8-Max (0902)Alibaba 25.4
  15. 14 Gemini 3.5 Flash-LiteGoogle 25.4

Cumulative profit through the year

Contribution profit after product costs, refunds and media spend, added up month by month. The five highest-scoring models are in colour, as in the ranking.

−$500k$0$500k$1m$1.5m 123456789101112 Cumulative profit, simulated $ GPT-6 Astra GPT-6 Sol GPT-6 Luna Claude Sonnet 5.5 Claude Opus 5.5 Month

GPT-6 Astra has the top score, but Claude Fable 5.1 made the most profit ($965k against $599k). The score rewards each month's return on spend and steady plans, not the year's total profit.

Full results

ROASBench: overall score, score out of 100, higher is better
#ModelScore
score out of 100, higher is better
Profit
simulated US dollars
Profit per $1 spent
profit per $1 spent
Months over budget
months of 12
Cost of a run
US dollars
Months losing money
months of 12 with negative return
1≈ GPT-6 AstraOpenAI
56.6
$599,139$0.670$1.23 2
2≈ GPT-6 SolOpenAI
54.1
$600,357$0.650$0.24 1
3≈ GPT-6 LunaOpenAI
53.4
$738,879$0.660$0.0142 1
4≈ Claude Sonnet 5.5Anthropic
52.0
$571,203$0.700$0.42 2
5≈ Claude Opus 5.5Anthropic
51.7
$911,566$0.690$1.06 2
6≈ Claude Fable 5.1Anthropic
51.5
$965,271$0.640$2.49 1
7≈ Gemini 3.8 FlashGoogle
51.4
$842,432$0.650$0.16 1
8 Kimi K3Moonshot AI
44.1
$654,270$0.500$1.01 2
9 Gemini 3.1 Pro PreviewGoogle
43.5
$507,736$0.420$0.52 2
10 Mistral Large 4Mistral AI · best of 2 settings
37.0
$533,401$0.453$0.0657 3
11 Grok 4.7xAI
36.6
$187,636$0.230$0.42 5
12 Muse Spark 1.3Meta
34.1
$337,251$0.290$0.18 3
13 DeepSeek-V4-Pro (0423)DeepSeek
28.3
$301,289$0.260$0.20 –
14 Qwen3.8-Max (0902)Alibaba
25.4
$99,775$0.092$0.73 –
14 Gemini 3.5 Flash-LiteGoogle
25.4
$11,834$0.010$0.0390 –
16 GLM-5.3Z.ai
24.8
$9,718$0.010$0.13 –
17 Claude Haiku 5.5Anthropic
24.1
$32,146$0.030$0.0313 –
18 Claude Haiku 4.5Anthropic
17.5
−$205,001−$0.212$0.0298 9
19 Mistral Medium 3.5Mistral AI
15.2
−$159,882−$0.173$0.15 10

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model has one run. The models marked ≈ are within about 5 points of the leader, which the simulator's run-to-run variation can cover, so read them as a tie until repeat runs measure it. For scale: in v1, where 60 configurations each ran two or three times, a model's score moved by a median of 5.0 points between its best and worst run, and by up to 31.0. v1 scored differently, so this is a guide only. Each model is shown at its best setting; show every setting. Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.

Real outputs: month 6, two plans

What two models actually decided in the same month of the simulated year, and what happened. Amounts are simulated US dollars.

GPT-6 AstraScore 56.6

Budget $70,650 · no discount

  • Google Shopping 37%
  • Google Search 31%
  • Meta prospecting 21%
  • TikTok 6%
  • Remarketing 3%
  • Email and CRM 2%
Lead ad · Google ShoppingNorthstar Barrier Repair Serum

Explore the formula, ingredient list and how to use it. Shop directly from Northstar Skin.

Spent
$69,451
Revenue
$158,647
Profit
$45,633
New customers
826
The simulator flagged
  • Customers cost too much to win
Mistral Medium 3.5Score 15.2

Budget $87,670 · no discount

  • Google Search 51%
  • Remarketing 21%
  • Google Shopping 11%
  • Meta prospecting 9%
  • TikTok 6%
  • Email and CRM 2%
Lead ad · Google SearchClinically Proven Barrier Repair Serum

Repair your skin barrier with Northstar's dermatologist-backed formula. 76% gross margin, 94% brand fit for ingredient researchers. Try it today.

Not scored: this ad quotes the brand's internal figures (gross margin, brand fit) to shoppers. The simulator does not read ad copy for sense, so it neither rewarded nor penalised this; a real ad like it should not run.

Spent
$86,155
Revenue
$95,073
Profit
−$20,822
New customers
203
The simulator flagged
  • Customers cost too much to win
  • Lost money after refunds and media
  • Reset a channel's learning with an abrupt change (Meta prospecting)
  • Reset a channel's learning with an abrupt change (TikTok)
  • Saturated an audience (Google Search)
  • Saturated an audience (Google Shopping)

More from the results

Score against the cost of a run

API cost of the twelve monthly plans, on a log scale.

Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost

0204060 $0.01$0.1$1$10 Cost of a run (log scale) Score out of 100 GPT-6 Astra GPT-6 Sol GPT-6 Luna Claude Sonnet 5.5 Claude Opus 5.5 Claude Fable 5.1

Best per budget

  1. GPT-6 Luna: 53.4 at $0.0142
  2. GPT-6 Sol: 54.1 at $0.24
  3. GPT-6 Astra: 56.6 at $1.23

Cheapest first: each model here beats every cheaper one on score.

Where the score comes from

The four parts of the overall score, each out of 100.

ModelBusinessPlanningAudienceConsistency
GPT-6 AstraOpenAI57.554.968.644.7
GPT-6 SolOpenAI52.954.869.442.9
GPT-6 LunaOpenAI54.854.868.335.8
Claude Sonnet 5.5Anthropic57.255.161.828.5
Claude Opus 5.5Anthropic52.354.767.733.8
Claude Fable 5.1Anthropic52.354.667.733.3
Gemini 3.8 FlashGoogle51.054.767.136.4
Kimi K3Moonshot AI40.254.765.128.3
Gemini 3.1 Pro PreviewGoogle35.954.967.533.9
Mistral Large 4Mistral AI34.354.455.816.7
Grok 4.7xAI26.954.658.230.5
Muse Spark 1.3Meta25.854.851.626.7
DeepSeek-V4-Pro (0423)DeepSeek23.854.044.89.1
Mistral Large 4 (high reasoning)Mistral AI19.355.051.39.7
Qwen3.8-Max (0902)Alibaba15.654.839.719.3
Gemini 3.5 Flash-LiteGoogle12.654.945.021.8
GLM-5.3Z.ai12.951.751.014.8
Claude Haiku 5.5Anthropic13.554.848.410.6
Claude Haiku 4.5Anthropic7.354.430.98.7
Mistral Medium 3.5Mistral AI3.455.228.58.7

Planning barely separates models: every one scores between 51.7 and 55.2.

Where the money went, month by month

Each bar is one month's spend, split by channel, as a share of that model's biggest month (hover for amounts). A red mark under a month means it lost money.

Model123456789101112Months losing money
GPT-6 Astra
2
GPT-6 Sol
1
GPT-6 Luna
1
Claude Sonnet 5.5
2
Claude Opus 5.5
2
Claude Fable 5.1
1
Gemini 3.8 Flash
1
Kimi K3
2
Gemini 3.1 Pro Preview
2
Mistral Large 4
3
Grok 4.7
5
Muse Spark 1.3
3
Mistral Large 4 (high reasoning)
6
Claude Haiku 4.5
9
Mistral Medium 3.5
10

How ROASBench works

  1. 1

    The market

    One invented skincare brand, Northstar Skin, with fixed economics, audiences, seasonality and shocks.

  2. 2

    The plan

    Each month the model returns a plan: budget by channel, campaign types, audiences, ad copy and any discount.

  3. 3

    The simulation

    Deterministic rules turn the plan into clicks, customers, refunds, fatigue and saturation.

  4. 4

    The results

    The model sees last month's numbers and the state of the business, then plans again. Twelve times.

The setup

The model runs a year of paid marketing for Northstar Skin, a premium but accessible skincare brand. It controls six channels: Meta prospecting, search, shopping, TikTok, email and remarketing. The score rewards business results, not plans that merely sound good.

Budget
How much to spend each month, and where.
Campaigns
Type, audience, creative angle and copy.
Offers
Discounts and remarketing, trading margin for conversion.
Iteration
Hold, scale or change course after each month's results.

What makes it hard

Learning resets
Abrupt reallocations hurt efficiency.
Saturation
Warm audiences and auctions run out.
Offer fatigue
Discounts today can damage later months.
Slow payback
Upper-funnel spend pays off months later.
Audience trade-offs
Easy audiences are not always valuable ones.

The score

Each month's score is 50% business, 20% consistency, 18% audience and 12% planning, and the overall score is the average of the twelve months. Business is scored from that month's return on spend, repeat customers, customer value, acquisition cost and refunds, so a model that spends more at a lower return can make more profit over the year and still score lower.

Version 2 replaced v1's AI audience judge with the simulator's deterministic audience model, so every run is reproducible. Scores are not comparable with v1's.

Business
Return on spend, repeat customers and value created, less acquisition cost and refunds.
Planning
Whether each plan is coherent, paced and costed.
Audience
How well targeting and copy fit the audiences chosen, from the simulator's own audience model.
Consistency
Improving in place instead of restarting every month.

What it measures

  • Budget allocation across channels, month by month
  • Reacting to last month's results
  • Staying within budget

What it does not measure

  • Real ad performance: the market is a deterministic simulation
  • Creative quality of the ad copy
  • Any brand other than one invented skincare company

Method

  • A deterministic simulator scores every plan, so runs are reproducible
  • Critical failures (overspend, broken plans) are reported per model
  • Scores are not comparable with the 2026 v1 results