Benchmarks / ROASBench
Measured by Spring Prompt
ROASBench
Given twelve months of paid-marketing decisions for a simulated brand, which models grow the business and which overspend?
- Results dated
- 29 Sep 2026
- Models
- 17
- Unit
- US dollars
- Licence
- Spring Prompt original
| # | Model | Cost of a run US dollars, lower is better | Overall score score out of 100 | Return on ad spend profit per £1 spent | Months over budget months of 12 |
|---|---|---|---|---|---|
| 1 | GPT-6 LunaOpenAI |
$0.0142
|
53.4 | £0.66 | 0 |
| 2 | Claude Haiku 4.5Anthropic |
$0.0298
|
17.5 | −£0.21 | 2 |
| 3 | Gemini 3.5 Flash LiteGoogle |
$0.0390
|
25.4 | £0.01 | 0 |
| 4 | GLM 5.3Z.ai |
$0.13
|
24.8 | £0.01 | 0 |
| 5 | Mistral Medium 3.5Mistral |
$0.15
|
15.2 | −£0.17 | 3 |
| 6 | Gemini 3.8 FlashGoogle |
$0.16
|
51.4 | £0.65 | 0 |
| 7 | Muse Spark 1.3Meta |
$0.18
|
34.1 | £0.29 | 0 |
| 8 | DeepSeek V4 Pro 0423DeepSeek |
$0.20
|
28.2 | £0.26 | 0 |
| 9 | GPT-6 SolOpenAI |
$0.24
|
54.1 | £0.65 | 0 |
| 10 | Claude Sonnet 5.5Anthropic |
$0.42
|
52.0 | £0.70 | 0 |
| 11 | Grok 4.7xAI |
$0.42
|
36.6 | £0.23 | 0 |
| 12 | Gemini 3.1 Pro PreviewGoogle |
$0.52
|
43.5 | £0.42 | 0 |
| 13 | Qwen3.8 Max (0902)Alibaba |
$0.73
|
25.4 | £0.09 | 2 |
| 14 | Kimi K3Moonshot AI |
$1.01
|
44.1 | £0.50 | 0 |
| 15 | Claude Opus 5.5Anthropic |
$1.06
|
51.6 | £0.69 | 0 |
| 16 | GPT-6 AstraOpenAI |
$1.23
|
56.6 | £0.67 | 0 |
| 17 | Claude Fable 5.1Anthropic |
$2.49
|
51.5 | £0.64 | 0 |
Each model runs at its provider's default reasoning setting. Some providers think at length by default and others barely at all, so this is what you get without tuning.
What it measures
- Budget allocation across channels, month by month
- Reacting to last month's results
- Staying within budget
What it does not measure
- Real ad performance: the market is a deterministic simulation
- Creative quality of the ad copy
- Any brand other than one invented skincare company
Method
- A deterministic simulator scores every plan, so runs are reproducible
- Critical failures (overspend, broken plans) are reported per model
- Scores are not comparable with the 2026 v1 results