Benchmarks / SurveyBench

Measured by Spring Prompt

SurveyBench

Given a research brief, which models plan a sound user survey, and then draw only the conclusions its responses support?

Last updated 8 Oct 2026

Results dated
8 Oct 2026
Models
8
Unit
% of answers
Licence
Spring Prompt original
Judge
GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.1 Pro, by majority · bias check
Runs
1 per task

Numbers right: Claude Haiku 4.5

8 results · % of answers, higher is better · ≈ cannot be told apart from the leader. Choose a model to highlight it.Clear highlight

  1. 1≈ Claude Haiku 5.5Anthropic 100.0%
  2. 1≈ Claude Opus 5.5Anthropic 100.0%
  3. 1≈ Gemini 3.1 Pro PreviewGoogle 100.0%
  4. 1≈ GPT-6.1 SolOpenAI 100.0%
  5. 5 Mistral Large 4Mistral AI 98.9%
  6. 6 Gemini 3.5 Flash-LiteGoogle 75.5%
  7. 6 Mistral Medium 3.5Mistral AI 75.5%
  8. 8 Claude Haiku 4.5Anthropic 63.7%

Analysis rating against cost

What one task (a survey plan and an analysis) cost at the run date's prices, on a log scale. Up and to the left is better.

Named: the six best and the best for the moneyOther models (hover for names)Best score at each cost

05001,0001,5002,000 $0.01$0.1$1 Cost per task, US dollars (log scale) Analysis rating GPT-6.1 Sol Claude Opus 5.5 Claude Haiku 5.5 Mistral Large 4 Gemini 3.1 Pro Preview Gemini 3.5 Flash-Lite

Best per budget

  1. Gemini 3.5 Flash-Lite: 662 at $0.0144
  2. Claude Haiku 5.5: 1,412 at $0.0185
  3. GPT-6.1 Sol: 1,650 at $0.14

Cheapest first: each model here beats every cheaper one on score.

Full results

SurveyBench: numbers right, % of answers, higher is better
#ModelNumbers right
% of answers, higher is better
Analysis rating
rating
Research score
points of 100
Sound research
% of tasks
Traps handled
% of traps
Cost per task
US dollars
1≈ Claude Haiku 5.5Anthropic
100.0%
1,41290.211.1%
1 of 9
94.2%$0.0185
1≈ Claude Opus 5.5Anthropic
100.0%
1,64696.155.6%
5 of 9
95.7%$0.46
1≈ Gemini 3.1 Pro PreviewGoogle
100.0%
81071.30.0%
0 of 9
71.0%$0.33
1≈ GPT-6.1 SolOpenAI
100.0%
1,65097.066.7%
6 of 9
95.7%$0.14
5 Mistral Large 4Mistral AI
98.9%
81571.20.0%
0 of 9
78.7%$0.0924
6 Gemini 3.5 Flash-LiteGoogle
75.5%
66255.60.0%
0 of 9
56.5%$0.0144
6 Mistral Medium 3.5Mistral AI
75.5%
54064.20.0%
0 of 9
59.4%$0.0730
8 Claude Haiku 4.5Anthropic
63.7%
46564.90.0%
0 of 9
62.3%$0.0635

Swipe the table sideways for more columns.

Ranks follow the score as shown, so equal numbers share a rank. Each model runs at its provider's default reasoning setting, once per task. The rating comes from head-to-head comparisons of the analyses, so it keeps separating models after several pass the checks; ranges are 95% intervals from resampling the comparisons. Models whose ranges overlap are not reliably different.

See the surveys and analyses

Every model's survey and analysis for every public task. Open a task to see the brief and the data every model analysed, or a model to read what it wrote, each number against the true value and what the judges flagged.

Why Fernway customers cancel · 2 of 8 sound

Whether to build one-tap skipping of any week (about four months of product work) or cut the price of the plans by 20%, to win back and keep customers. The brief and the data →

Why Tallyroom trials don't convert · 1 of 8 sound

Where to put next quarter's product and sales effort to lift trial-to-paid conversion: a new accounting integration, assisted setup, a cheaper plan, or a shorter trial. The brief and the data →

Why Pennywell customers don't use savings pots · 2 of 8 sound

Whether to run an awareness and onboarding push for savings pots (in-app prompts and a set-up nudge on payday), to raise the pots rate to 3.85% AER, or both. The brief and the data →

Why Quaymark's customer score fell · 2 of 8 sound

Whether to reverse the per-shipment fee and bring back the subscription, invest in shipment tracking and proactive delay alerts, or add account managers for mid-sized customers. The brief and the data →

What Hollinbrook residents think of three-weekly bin collections · 1 of 8 sound

Whether to move general waste collection to every three weeks across the whole district, to do so with extra support for households that need it (a larger bin, or extra collections of nappies and hygiene waste), to pilot it in some areas first, or to drop the proposal. The brief and the data →

What Ledgerline's accountancy practices want built next · 2 of 8 sound

Which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

How Ashgrove patients find online-first booking · 0 of 8 sound

Whether to keep online-first booking as it is, restore a morning phone booking line for every patient, or add a dedicated phone line and in-person help at reception only for patients who need them (over-75s, patients with a disability, and patients with limited English). The brief and the data →

Should Corefit gyms stop opening 24 hours? · 0 of 8 sound

Whether to close all 14 gyms between 11pm and 5am, close overnight only at the quietest gyms, or keep every gym open 24 hours. The brief and the data →

Why Riverside Food Network is losing volunteers · 2 of 8 sound

How to spend next year's £45,000 for volunteer support: reimburse volunteers' travel costs, fund a part-time volunteer coordinator with online shift booking, or do both at a smaller scale. The brief and the data →

More from the results

Right numbers are not the same as a sound reading

Share of the set numeric questions answered correctly, against the share of planted traps in the data the analysis handled. Several models get the arithmetic right and still miss what the data cannot show.

Numbers rightTraps handled
Claude Haiku 5.5
100.0% · 94.2%
Claude Opus 5.5
100.0% · 95.7%
Gemini 3.1 Pro Preview
100.0% · 71.0%
GPT-6.1 Sol
100.0% · 95.7%
Mistral Large 4
98.9% · 78.7%
Gemini 3.5 Flash-Lite
75.5% · 56.5%
Mistral Medium 3.5
75.5% · 59.4%
Claude Haiku 4.5
63.7% · 62.3%

Why the research is not sound

Share of tasks failing each check. A task can fail several at once; any one means the survey or the analysis needs fixing before it is used.

ModelTraps missedWrong numbersUnsupported findingsMistaken next stepsWrong recommendationLeading or double-barrelledSurvey structureObjectives unanswerableStakeholder's view presumedPrecision ignored
GPT-6.1 SolOpenAI33.3%0.0%0.0%11.1%11.1%0.0%11.1%0.0%0.0%0.0%
Claude Opus 5.5Anthropic33.3%0.0%22.2%0.0%0.0%0.0%11.1%0.0%0.0%0.0%
Claude Haiku 5.5Anthropic22.2%0.0%77.8%22.2%11.1%33.3%22.2%0.0%0.0%0.0%
Mistral Large 4Mistral AI100.0%22.2%100.0%66.7%44.4%33.3%11.1%11.1%22.2%55.6%
Gemini 3.1 Pro PreviewGoogle88.9%0.0%77.8%33.3%44.4%44.4%22.2%0.0%22.2%55.6%
Gemini 3.5 Flash-LiteGoogle100.0%100.0%66.7%44.4%22.2%77.8%55.6%33.3%55.6%100.0%
Mistral Medium 3.5Mistral AI100.0%100.0%100.0%77.8%44.4%55.6%11.1%11.1%33.3%100.0%
Claude Haiku 4.5Anthropic88.9%100.0%100.0%55.6%44.4%66.7%0.0%22.2%22.2%100.0%

How SurveyBench works

  1. 1

    The brief

    An invented organisation's research brief: the decision, four objectives, who to survey, limits on length and sample, and a stakeholder's firm view.

  2. 2

    The survey

    The model writes the survey as it would be sent: screener, questions, display logic, sample plan and analysis plan.

  3. 3

    The analysis

    Every model then analyses the same fielded survey: crosstabs with bases, fielding notes and open-text answers from simulated respondents.

  4. 4

    Head to head

    On each task the judges compare analyses in pairs and pick the one they would rather send to the decision-makers. The votes become the rating.

The task

Survey data rarely says what it seems to at first glance. The risk is not an ugly chart: it is a confident finding from twelve respondents, a satisfaction score asked only of the people who got through, or a stakeholder's hunch written into the question.

Each task is an invented organisation (a meal-kit subscription, a bank, software firms, a council, a GP practice, gyms, a charity) with a brief written for the benchmark. The responses are simulated from respondent-level data with planted effects, so every true number, and every way of being misled, is known before any model sees it.

What makes the research sound

The survey
Within the brief's limits on questions and minutes; every objective answerable; no leading or double-barrelled question; the stakeholder's view tested neutrally; and a sample plan that says plainly when the precision asked for cannot be met.
The numbers
Ten to twelve set questions per task: shares, weighted figures, margins of error, significance tests, differences and counts of open-text answers, each checked against the true value.
The traps
Skewed samples, small bases, questions asked of only some respondents, averages pulled by a few extreme answers, trends broken by a change of method, stated intentions, organised campaigns and reversals within subgroups.
The advice
A recommendation that follows from the data, next steps that would settle what the survey cannot, and none of the mistakes a careless reading leads to.

Why a rating as well

Sound research is a bar: once the best models clear it on most tasks, it stops telling them apart. The head-to-head rating does not run out of room, because an analysis only has to be preferred to another one. Each analysis meets about six others per task; a 400-point gap means the higher-rated model's analysis is preferred about ten times out of eleven.

What it measures

  • Survey questions: leading, double-barrelled, assumptive or ambiguous
  • Whether the survey can answer the brief within its limits
  • Numbers read from crosstabs, weighted, and tested for significance
  • Whether the analysis is misled by weaknesses planted in the data
  • Next steps, and what would change the recommendation

What it does not measure

  • Your customers, your data or your survey tool
  • Interviews, usability tests or other qualitative research
  • Fielding the survey or cleaning raw response files

Method

  • Respondents are simulated, so every true number is known exactly
  • Numbers and survey structure are checked by code
  • Three judges from three vendors decide each judged check by majority
  • No judge compares analyses from its own vendor
  • Three tasks are held back so the set can be refreshed

Checking the judge

Every judged check is decided by a majority of three judges: GPT-6.1 Sol, Claude Opus 5.5 and Gemini 3.1 Pro. On 12 tasks and seven models, corrected for chance (Fleiss' kappa), they agreed most on next steps (0.81 to 0.85), traps (0.73) and whether a survey faces the precision it was asked for (0.93), and least on double-barrelled questions (0.38), whether every objective can be answered (0.39) and which findings are unsupported (0.49). No judge favoured its own vendor's model. In head-to-head comparisons a judge never votes on a pair that includes its own vendor's analysis. A sample of human ratings is still to come.