16 models had none: Claude Fable 5.1, Claude Haiku 5.5, Claude Opus 5.5, Claude Sonnet 5.5, DeepSeek-V4-Pro (0423), GLM-5.3, GPT-6 Astra, GPT-6 Luna, GPT-6 Sol, Gemini 3.1 Pro Preview, Gemini 3.5 Flash-Lite, Gemini 3.8 Flash, Grok 4.7, Kimi K3, Mistral Large 4 (high reasoning), Muse Spark 1.3.
1.522.533.5
Cumulative profit through the year
Contribution profit after product costs, refunds and media spend, added up month by month. The five highest-scoring models are in colour, as in the ranking.
GPT-6 Astra has the top score, but Claude Fable 5.1 made the most profit ($965k against $599k). The score rewards each month's return on spend and steady plans, not the year's total profit.
Full results
ROASBench: months over budget, months of 12, lower is better
#
Model
Months over budget months of 12, lower is better
Score score out of 100
Profit simulated US dollars
Profit per $1 spent profit per $1 spent
Cost of a run US dollars
Months losing money months of 12 with negative return
Ranks follow the score as shown, so equal numbers share a rank. Each model has one run. The models marked ≈ are within about 5 points of the leader, which the simulator's run-to-run variation can cover, so read them as a tie until repeat runs measure it. For scale: in v1, where 60 configurations each ran two or three times, a model's score moved by a median of 5.0 points between its best and worst run, and by up to 31.0. v1 scored differently, so this is a guide only. Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.
Real outputs: month 6, two plans
What two models actually decided in the same month of the simulated year, and what happened. Amounts are simulated US dollars.
GPT-6 AstraScore 56.6
Budget $70,650 · no discount
Google Shopping 37%
Google Search 31%
Meta prospecting 21%
TikTok 6%
Remarketing 3%
Email and CRM 2%
Lead ad · Google ShoppingNorthstar Barrier Repair Serum
Explore the formula, ingredient list and how to use it. Shop directly from Northstar Skin.
Spent
$69,451
Revenue
$158,647
Profit
$45,633
New customers
826
The simulator flagged
Customers cost too much to win
Mistral Medium 3.5Score 15.2
Budget $87,670 · no discount
Google Search 51%
Remarketing 21%
Google Shopping 11%
Meta prospecting 9%
TikTok 6%
Email and CRM 2%
Lead ad · Google SearchClinically Proven Barrier Repair Serum
Repair your skin barrier with Northstar's dermatologist-backed formula. 76% gross margin, 94% brand fit for ingredient researchers. Try it today.
Not scored: this ad quotes the brand's internal figures (gross margin, brand fit) to shoppers. The simulator does not read ad copy for sense, so it neither rewarded nor penalised this; a real ad like it should not run.
Spent
$86,155
Revenue
$95,073
Profit
−$20,822
New customers
203
The simulator flagged
Customers cost too much to win
Lost money after refunds and media
Reset a channel's learning with an abrupt change (Meta prospecting)
Reset a channel's learning with an abrupt change (TikTok)
Saturated an audience (Google Search)
Saturated an audience (Google Shopping)
More from the results
Score against the cost of a run
API cost of the twelve monthly plans, on a log scale.
Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost
Best per budget
GPT-6 Luna: 53.4 at $0.0142
GPT-6 Sol: 54.1 at $0.24
GPT-6 Astra: 56.6 at $1.23
Cheapest first: each model here beats every cheaper one on score.
Where the score comes from
The four parts of the overall score, each out of 100.
Model
Business
Planning
Audience
Consistency
GPT-6 AstraOpenAI
57.5
54.9
68.6
44.7
GPT-6 SolOpenAI
52.9
54.8
69.4
42.9
GPT-6 LunaOpenAI
54.8
54.8
68.3
35.8
Claude Sonnet 5.5Anthropic
57.2
55.1
61.8
28.5
Claude Opus 5.5Anthropic
52.3
54.7
67.7
33.8
Claude Fable 5.1Anthropic
52.3
54.6
67.7
33.3
Gemini 3.8 FlashGoogle
51.0
54.7
67.1
36.4
Kimi K3Moonshot AI
40.2
54.7
65.1
28.3
Gemini 3.1 Pro PreviewGoogle
35.9
54.9
67.5
33.9
Mistral Large 4Mistral AI
34.3
54.4
55.8
16.7
Grok 4.7xAI
26.9
54.6
58.2
30.5
Muse Spark 1.3Meta
25.8
54.8
51.6
26.7
DeepSeek-V4-Pro (0423)DeepSeek
23.8
54.0
44.8
9.1
Mistral Large 4 (high reasoning)Mistral AI
19.3
55.0
51.3
9.7
Qwen3.8-Max (0902)Alibaba
15.6
54.8
39.7
19.3
Gemini 3.5 Flash-LiteGoogle
12.6
54.9
45.0
21.8
GLM-5.3Z.ai
12.9
51.7
51.0
14.8
Claude Haiku 5.5Anthropic
13.5
54.8
48.4
10.6
Claude Haiku 4.5Anthropic
7.3
54.4
30.9
8.7
Mistral Medium 3.5Mistral AI
3.4
55.2
28.5
8.7
Planning barely separates models: every one scores between 51.7 and 55.2.
Where the money went, month by month
Each bar is one month's spend, split by channel, as a share of that model's biggest month (hover for amounts). A red mark under a month means it lost money.
Search
Shopping
Meta
TikTok
Email
Remarketing
Lost money that month
Model
1
2
3
4
5
6
7
8
9
10
11
12
Months losing money
GPT-6 Astra
2
GPT-6 Sol
1
GPT-6 Luna
1
Claude Sonnet 5.5
2
Claude Opus 5.5
2
Claude Fable 5.1
1
Gemini 3.8 Flash
1
Kimi K3
2
Gemini 3.1 Pro Preview
2
Mistral Large 4
3
Grok 4.7
5
Muse Spark 1.3
3
Mistral Large 4 (high reasoning)
6
Claude Haiku 4.5
9
Mistral Medium 3.5
10
How ROASBench works
1
The market
One invented skincare brand, Northstar Skin, with fixed economics, audiences, seasonality and shocks.
2
The plan
Each month the model returns a plan: budget by channel, campaign types, audiences, ad copy and any discount.
3
The simulation
Deterministic rules turn the plan into clicks, customers, refunds, fatigue and saturation.
4
The results
The model sees last month's numbers and the state of the business, then plans again. Twelve times.
The setup
The model runs a year of paid marketing for Northstar Skin, a premium but accessible skincare brand. It controls six channels: Meta prospecting, search, shopping, TikTok, email and remarketing. The score rewards business results, not plans that merely sound good.
Budget
How much to spend each month, and where.
Campaigns
Type, audience, creative angle and copy.
Offers
Discounts and remarketing, trading margin for conversion.
Iteration
Hold, scale or change course after each month's results.
What makes it hard
Learning resets
Abrupt reallocations hurt efficiency.
Saturation
Warm audiences and auctions run out.
Offer fatigue
Discounts today can damage later months.
Slow payback
Upper-funnel spend pays off months later.
Audience trade-offs
Easy audiences are not always valuable ones.
The score
Each month's score is 50% business, 20% consistency, 18% audience and 12% planning, and the overall score is the average of the twelve months. Business is scored from that month's return on spend, repeat customers, customer value, acquisition cost and refunds, so a model that spends more at a lower return can make more profit over the year and still score lower.
Version 2 replaced v1's AI audience judge with the simulator's deterministic audience model, so every run is reproducible. Scores are not comparable with v1's.
Business
Return on spend, repeat customers and value created, less acquisition cost and refunds.
Planning
Whether each plan is coherent, paced and costed.
Audience
How well targeting and copy fit the audiences chosen, from the simulator's own audience model.
Consistency
Improving in place instead of restarting every month.
What it measures
Budget allocation across channels, month by month
Reacting to last month's results
Staying within budget
What it does not measure
Real ad performance: the market is a deterministic simulation
Creative quality of the ad copy
Any brand other than one invented skincare company
Method
A deterministic simulator scores every plan, so runs are reproducible
Critical failures (overspend, broken plans) are reported per model
Scores are not comparable with the 2026 v1 results