Benchmarks / ROASBench

Measured by Spring Prompt

ROASBench

Given twelve months of paid-marketing decisions for a simulated brand, which models grow the business and which overspend?

Results dated
29 Sep 2026
Models
17
Unit
months of 12
Licence
Spring Prompt original

Cumulative profit through the year

Contribution profit after product costs, refunds and media spend, added up month by month. The five most profitable models are in colour.

−$500k$0$500k$1,000k$1,500k 123456789101112 Cumulative profit, simulated $ GPT-6 Sol: $600,357 by month 12 GPT-6 Astra: $599,139 by month 12 Claude Sonnet 5.5: $571,203 by month 12 Gemini 3.1 Pro Preview: $507,736 by month 12 Muse Spark 1.3: $337,251 by month 12 DeepSeek V4 Pro 0423: $301,289 by month 12 Grok 4.7: $187,636 by month 12 Qwen3.8 Max (0902): $99,775 by month 12 Gemini 3.5 Flash Lite: $11,834 by month 12 GLM 5.3: $9,718 by month 12 Mistral Medium 3.5: $-159,882 by month 12 Claude Haiku 4.5: $-205,001 by month 12 Claude Fable 5.1: $965,271 by month 12 Claude Fable 5.1 Claude Opus 5.5: $911,566 by month 12 Claude Opus 5.5 Gemini 3.8 Flash: $842,432 by month 12 Gemini 3.8 Flash GPT-6 Luna: $738,878 by month 12 GPT-6 Luna Kimi K3: $654,270 by month 12 Kimi K3 Month

Full results

ROASBench: months over budget, months of 12, lower is better
#ModelROASBench: months over budget
months of 12, lower is better
Score
score out of 100
Profit
simulated US dollars
Return per $1
profit per $1 spent
Cost of a run
US dollars
1 Claude Fable 5.1Anthropic
0
51.5$965,271$0.64$2.49
1 Claude Opus 5.5Anthropic
0
51.6$911,566$0.69$1.06
1 Claude Sonnet 5.5Anthropic
0
52.0$571,203$0.70$0.42
1 DeepSeek V4 Pro 0423DeepSeek
0
28.2$301,289$0.26$0.20
1 Gemini 3.1 Pro PreviewGoogle
0
43.5$507,736$0.42$0.52
1 Gemini 3.5 Flash LiteGoogle
0
25.4$11,834$0.01$0.0390
1 Gemini 3.8 FlashGoogle
0
51.4$842,432$0.65$0.16
1 Muse Spark 1.3Meta
0
34.1$337,251$0.29$0.18
1 Kimi K3Moonshot AI
0
44.1$654,270$0.50$1.01
1 GPT-6 AstraOpenAI
0
56.6$599,139$0.67$1.23
1 GPT-6 LunaOpenAI
0
53.4$738,878$0.66$0.0142
1 GPT-6 SolOpenAI
0
54.1$600,357$0.65$0.24
1 Grok 4.7xAI
0
36.6$187,636$0.23$0.42
1 GLM 5.3Z.ai
0
24.8$9,718$0.01$0.13
15 Qwen3.8 Max (0902)Alibaba
2
25.4$99,775$0.09$0.73
15 Claude Haiku 4.5Anthropic
2
17.5−$205,001−$0.21$0.0298
17 Mistral Medium 3.5Mistral
3
15.2−$159,882−$0.17$0.15

Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.

Real outputs: month 6, two plans

What two models actually decided in the same month of the simulated year, and what happened. Amounts are simulated US dollars.

GPT-6 AstraScore 56.6

Budget $70,650 · no discount

  • Google Shopping 37%
  • Google Search 31%
  • Meta prospecting 21%
  • TikTok 6%
  • Remarketing 3%
  • Email and CRM 2%
Lead ad · Google ShoppingNorthstar Barrier Repair Serum

Explore the formula, ingredient list and how to use it. Shop directly from Northstar Skin.

Spent
$69,451
Revenue
$158,647
Profit
$45,633
New customers
826
The simulator flagged
  • Customers cost too much to win
Mistral Medium 3.5Score 15.2

Budget $87,670 · no discount

  • Google Search 51%
  • Remarketing 21%
  • Google Shopping 11%
  • Meta prospecting 9%
  • TikTok 6%
  • Email and CRM 2%
Lead ad · Google SearchClinically Proven Barrier Repair Serum

Repair your skin barrier with Northstar's dermatologist-backed formula. 76% gross margin, 94% brand fit for ingredient researchers. Try it today.

Spent
$86,155
Revenue
$95,073
Profit
−$20,822
New customers
203
The simulator flagged
  • Customers cost too much to win
  • Lost money after refunds and media
  • Reset a channel's learning with an abrupt change (Meta prospecting)
  • Reset a channel's learning with an abrupt change (TikTok)
  • Saturated an audience (Google Search)
  • Saturated an audience (Google Shopping)

More from the results

Score against the cost of a run

API cost of the twelve monthly plans, on a log scale.

0.020.040.060.0 $0.01$0.1$1$10 Cost of a run (log scale) Score out of 100 Gemini 3.8 Flash: 51.4, $0.16 Kimi K3: 44.1, $1.01 Gemini 3.1 Pro Preview: 43.5, $0.52 Grok 4.7: 36.6, $0.42 Muse Spark 1.3: 34.1, $0.18 DeepSeek V4 Pro 0423: 28.2, $0.20 Qwen3.8 Max (0902): 25.4, $0.73 Gemini 3.5 Flash Lite: 25.4, $0.0390 GLM 5.3: 24.8, $0.13 Claude Haiku 4.5: 17.5, $0.0298 Mistral Medium 3.5: 15.2, $0.15 GPT-6 Astra: 56.6, $1.23 GPT-6 Sol: 54.1, $0.24 GPT-6 Luna: 53.4, $0.0142 Claude Sonnet 5.5: 52.0, $0.42 Claude Opus 5.5: 51.6, $1.06 Claude Fable 5.1: 51.5, $2.49 GPT-6 Astra GPT-6 Sol GPT-6 Luna Claude Sonnet 5.5 Claude Opus 5.5 Claude Fable 5.1

Where the score comes from

The four parts of the overall score, each out of 100.

ModelBusinessPlanningAudienceConsistency
GPT-6 AstraOpenAI57.554.968.644.7
GPT-6 SolOpenAI52.954.869.442.9
GPT-6 LunaOpenAI54.854.868.235.8
Claude Sonnet 5.5Anthropic57.255.161.828.5
Claude Opus 5.5Anthropic52.354.667.733.8
Claude Fable 5.1Anthropic52.354.667.733.3
Gemini 3.8 FlashGoogle51.054.767.136.4
Kimi K3Moonshot AI40.254.765.128.3
Gemini 3.1 Pro PreviewGoogle35.954.967.533.9
Grok 4.7xAI26.954.658.230.5
Muse Spark 1.3Meta25.854.851.626.7
DeepSeek V4 Pro 0423DeepSeek23.854.044.89.1
Qwen3.8 Max (0902)Alibaba15.654.839.719.2
Gemini 3.5 Flash LiteGoogle12.654.945.021.8
GLM 5.3Z.ai12.951.751.014.8
Claude Haiku 4.5Anthropic7.354.430.98.7
Mistral Medium 3.5Mistral3.455.228.58.7

How ROASBench works

  1. 1

    The market

    One invented skincare brand, Northstar Skin, with fixed economics, audiences, seasonality and shocks.

  2. 2

    The plan

    Each month the model returns a plan: budget by channel, campaign types, audiences, ad copy and any discount.

  3. 3

    The simulation

    Deterministic rules turn the plan into clicks, customers, refunds, fatigue and saturation.

  4. 4

    The results

    The model sees last month's numbers and the state of the business, then plans again. Twelve times.

The setup

The model runs a year of paid marketing for Northstar Skin, a premium but accessible skincare brand. It controls six channels: Meta prospecting, search, shopping, TikTok, email and remarketing. The score rewards business results, not plans that merely sound good.

Budget
How much to spend each month, and where.
Campaigns
Type, audience, creative angle and copy.
Offers
Discounts and remarketing, trading margin for conversion.
Iteration
Hold, scale or change course after each month's results.

What makes it hard

Learning resets
Abrupt reallocations hurt efficiency.
Saturation
Warm audiences and auctions run out.
Offer fatigue
Discounts today can damage later months.
Slow payback
Upper-funnel spend pays off months later.
Audience trade-offs
Easy audiences are not always valuable ones.

The score

Version 2 replaced v1's AI audience judge with the simulator's deterministic audience model, so every run is reproducible. Scores are not comparable with v1's.

Business
Return on spend, repeat customers and value created, less acquisition cost and refunds.
Planning
Whether each plan is coherent, paced and costed.
Audience
How well targeting and copy fit the audiences chosen, from the simulator's own audience model.
Consistency
Improving in place instead of restarting every month.

What it measures

  • Budget allocation across channels, month by month
  • Reacting to last month's results
  • Staying within budget

What it does not measure

  • Real ad performance: the market is a deterministic simulation
  • Creative quality of the ad copy
  • Any brand other than one invented skincare company

Method

  • A deterministic simulator scores every plan, so runs are reproducible
  • Critical failures (overspend, broken plans) are reported per model
  • Scores are not comparable with the 2026 v1 results