Benchmarks / ROASBench

Measured by Spring Prompt

ROASBench

Given twelve months of paid-marketing decisions for a simulated brand, which models grow the business and which overspend?

Results dated
29 Sep 2026
Models
17
Unit
US dollars
Licence
Spring Prompt original
ROASBench: cost of a run, US dollars, lower is better
#ModelCost of a run
US dollars, lower is better
Overall score
score out of 100
Return on ad spend
profit per £1 spent
Months over budget
months of 12
1 GPT-6 LunaOpenAI
$0.0142
53.4£0.660
2 Claude Haiku 4.5Anthropic
$0.0298
17.5−£0.212
3 Gemini 3.5 Flash LiteGoogle
$0.0390
25.4£0.010
4 GLM 5.3Z.ai
$0.13
24.8£0.010
5 Mistral Medium 3.5Mistral
$0.15
15.2−£0.173
6 Gemini 3.8 FlashGoogle
$0.16
51.4£0.650
7 Muse Spark 1.3Meta
$0.18
34.1£0.290
8 DeepSeek V4 Pro 0423DeepSeek
$0.20
28.2£0.260
9 GPT-6 SolOpenAI
$0.24
54.1£0.650
10 Claude Sonnet 5.5Anthropic
$0.42
52.0£0.700
11 Grok 4.7xAI
$0.42
36.6£0.230
12 Gemini 3.1 Pro PreviewGoogle
$0.52
43.5£0.420
13 Qwen3.8 Max (0902)Alibaba
$0.73
25.4£0.092
14 Kimi K3Moonshot AI
$1.01
44.1£0.500
15 Claude Opus 5.5Anthropic
$1.06
51.6£0.690
16 GPT-6 AstraOpenAI
$1.23
56.6£0.670
17 Claude Fable 5.1Anthropic
$2.49
51.5£0.640

Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.

What it measures

  • Budget allocation across channels, month by month
  • Reacting to last month's results
  • Staying within budget

What it does not measure

  • Real ad performance: the market is a deterministic simulation
  • Creative quality of the ad copy
  • Any brand other than one invented skincare company

Method

  • A deterministic simulator scores every plan, so runs are reproducible
  • Critical failures (overspend, broken plans) are reported per model
  • Scores are not comparable with the 2026 v1 results