Benchmarks / SurveyBench / Why Tallyroom trials don't convert / GPT-6.1 Sol
GPT-6.1 Sol: Why Tallyroom trials don't convert
The decision: where to put next quarter's product and sales effort to lift trial-to-paid conversion: a new accounting integration, assisted setup, a cheaper plan, or a shorter trial. The brief and the data →
- Verdict
- ✓ Sound research
- Research score
- 97 of 100
- Analysis rating
- 1,650
- Head to head, this task
- won 8 of 9
Scorecard
Numbers 11 of 11 right
| Question | Its answer | True | |
|---|---|---|---|
| What percentage of respondents said they would need a Sage integration before paying? | ✓ | 23.7% | 23.7% |
| Weighted to the company-size mix of all lapsed trials, what percentage would need a Sage integration? | ✓ | 18.1% | 18.2% |
| Among businesses with 1-9 employees, what percentage gave 'didn't have time to set it up' as the main reason? | ✓ | 27.4% | 27.4% |
| How many of the open-text answers mention not having time to set Tallyroom up, or setup taking too long? | ✓ | 24 | 24 |
| What is the 95% margin of error, in percentage points, on the unweighted share of respondents needing a Sage integration? | ✓ | 3.9 points | 3.89 points |
| Is the difference in the share giving 'too expensive' as their main reason between businesses with 1-9 employees and those with 50-249 statistically significant at the 95% level? (1 for yes, 0 for no) | ✓ | No | No |
| Weighted to the company-size mix of all 1,640 lapsed trials, about how many teams' main reason was not having time to set it up? | ✓ | 355 | 355 |
| How many of the open-text answers from businesses with 1-9 employees mention not having time to set Tallyroom up, or setup taking too long? | ✓ | 14 | 14 |
| Weighted to the company-size mix of all lapsed trials, what percentage would give 'no integration with our accounting software' as their main reason? | ✓ | 16.3% | 16.3% |
| What percentage of the trial admins invited completed the survey? | ✓ | 28% | 28% |
| Among respondents whose main reason was no integration with their accounting software, what percentage said they would be likely (fairly or very) to pay if Tallyroom connected to it? | ✓ | 61.3% | 61.3% |
Traps in the data 7 of 7 handled
- ✓Teams on 'Other' accounting software (base 17) mostly cite missing integration.
- ✓Teams of 50-249 people are 22% of respondents but 10% of lapsed trials. They cite missing integration and use Sage far more, so unweighted figures overstate integration and Sage demand and understate setup time: unweighted, missing integration is the top reason; weighted, setup time is.
- ✓Q2 is multi-select: percentages are of respondents and add up to more than 100%.
- ✓Likelihood to pay is a hypothetical stated intention.
- ✓The head of sales believes the trial is too long and teams forget about it.
- ✓460 of 1,640 invited admins completed (about 28%), with a £10 voucher; those who respond may differ from those who do not.
- ✓Only teams that did not pay were surveyed, so the survey cannot show what distinguishes them from teams that did; paying teams may also have struggled with setup.
The analysis
Put the main near-term product and sales effort into a randomised assisted-setup pilot for businesses using supported accounting software, while retaining the 30-day trial. Combine guided configuration with invitations, approval setup and a first successfully exported expense, rather than merely adding reminders. In parallel, give Sage a bounded discovery and commercial-validation workstream: identify the exact Sage products used, technical scope, other blockers and prospects willing to make a real paid commitment. Do not commit the quarter to a generic Sage build from these tables alone. Test pricing structure rather than launching a blanket discount, and triage mobile and approval failures that would undermine either onboarding or integration. Audit the source data immediately; the ranking is provisional, but a controlled setup pilot remains a lower-commitment way to obtain causal evidence.
Shift more effort to Sage if audited trial records show sufficient demand for a clearly defined Sage product, prospects make credible paid commitments, and expected incremental retained revenue exceeds build and support costs relative to assisted setup. Shift to pricing if a controlled offer test increases retained contribution margin rather than merely subsidising teams that would already pay. Consider a 14-day trial if a randomised duration test improves paid conversion and retained revenue over a common follow-up window without reducing activation or increasing setup-related losses. Deprioritise assisted setup if its experiment produces little incremental conversion or costs more than the resulting customer contribution. A corrected size distribution or respondent linkage that materially changes the barrier ranking would also reopen prioritisation.
Findings
-
Prioritise an assisted-setup experiment next quarter. Setup is the largest of the four proposed barriers after adjusting for the population's company-size mix, and it can be addressed without first building another integration.
Setup was the main reason for 86/460 respondents (18.7%), rising to approximately 21.6% after size weighting, versus 16.3% for missing integrations and 16.6% for price. Applying the weighted setup share to 1,640 lapsed trials gives approximately 355 teams, not 355 recoverable customers. Among 1–9 employee respondents, 46/168 (27.4%) cited setup. The supplied open texts contain 24 setup-time mentions, including 14 from the labelled 1–9 employee segment.
-
There is substantial stated demand for Sage, but it supports targeted discovery and commercial validation rather than an immediate commitment to an unspecified 'Sage integration'.
109/460 respondents (23.7%) selected Sage integration as needed before paying; the size-weighted estimate is 18.1%. Among Sage users, 100/117 (85.5%) selected it, and 68/117 (58.1%) identified missing integration as their main reason for not paying. The nominal 95% margin of error on the unweighted 23.7% is approximately ±3.9 percentage points. Open texts name Sage 50, Sage 200, Sage Business Cloud and Sage Accounting, so the demand spans products whose integration requirements must be verified.
-
Do not cut the standard trial from 30 to 14 days on this evidence. Forgetting is a minority explanation, while insufficient setup time is much more prominent; shortening could worsen that barrier.
Only 1/460 respondents (0.2%) chose 'trial was too long'; 17/460 (3.7%) chose forgetting, compared with 86/460 (18.7%) choosing setup time and 7/460 (1.5%) choosing 'trial was too short'. The 24 setup-time texts describe interrupted work and unfinished configuration. These retrospective answers do not establish the causal effect of trial duration.
-
The finance director cannot use this survey to forecast how many teams will pay after integration. It identifies promising prospects, not a conversion rate.
Among the 93 respondents citing missing integration as their main reason, 57 (61.3%) said they were fairly or very likely to pay £6 per user per month; only 24 (25.8%) were very likely. Across all 460 respondents, stated likelihood was 37.6%. The tables do not provide Sage-specific payment likelihood, joint requirements or actual purchasing behaviour. The 61.3% must not be multiplied by all lapsed trials or all Sage users as a revenue forecast.
-
Investigate pricing structure before introducing a blanket cheaper plan. The texts point particularly to paying for occasional users, rather than proving that a lower per-user price is the best remedy.
73/460 respondents (15.9%) selected price as their main reason; 94/460 (20.4%) requested per-company pricing. There are 23 supplied texts explicitly discussing price or charging structure. The price-main-reason difference between 1–9 employee businesses (32/168, 19.0%) and 50–249 employee businesses (17/110, 15.5%) is not statistically significant using a conventional two-sided two-proportion test (p approximately 0.44). This does not prove equal sensitivity.
-
An integration or onboarding improvement will not resolve every blocker. Validate approval flexibility and mobile reliability alongside the conversion work.
58/460 (12.6%) selected approval workflow as their main reason and 101/460 (22.0%) requested custom approval chains. The 15 approval-related texts include batch approval, external approvers, multiple directors and expense-dependent routing. Mobile problems were the main reason for 29/460 (6.3%); 10 texts describe scanning or upload failures, with three also mentioning setup time.
-
Sage has the stronger unweighted integration-demand signal than FreeAgent, but the survey does not establish which integration would deliver the better return.
Sage was requested by 109/460 (23.7%), versus FreeAgent by 33/460 (7.2%). Within their respective user groups, 100/117 Sage users (85.5%) and 24/33 FreeAgent users (72.7%) requested the relevant integration. Missing integration was the main reason for 68/117 Sage users (58.1%) and 11/33 FreeAgent users (33.3%). Build cost, software-version coverage, addressable trial volumes and realised conversion are unknown.
-
Treat the size-weighted prioritisation as provisional until the respondent records are audited. The supplied data contain material inconsistencies that could change segment estimates.
The respondent size mix is 36.5%/39.6%/23.9%, versus 55%/35%/10% in the lapsed-trial population. Several text labels conflict with stated headcount: T010 is labelled 50–249 but says nine people; T023 is labelled 1–9 but says 15 employees; T164 is labelled 50–249 but says four people; T449 is labelled 1–9 but says 18. Also, Q2 reports zero 'didn't need it' responses in the 50–249 group, while several texts labelled in that group explicitly express that reason.
Caveats
- 460 of 1,640 invited admins completed the survey: 28.0%. This is a self-selected, incentivised respondent group, not a random sample. Nonrespondents may have different barriers; the £10 voucher does not remove that risk.
- The survey includes only lapsed, non-paying trials. It cannot directly estimate overall trial-to-paid conversion, identify differences from successful trials or establish which intervention causes payment.
- Company-size weighting uses the supplied population mix and rounded table percentages. It corrects the observed size imbalance only, not nonresponse within size groups, accounting-software imbalance or other selection effects.
- The approximately 355 setup-blocked teams is an extrapolation conditional on respondents representing nonrespondents within company-size groups. It is not an observed population count or a forecast of incremental customers.
- The ±3.9 percentage-point margin is a conventional unweighted binomial margin, without finite-population correction. It excludes selection bias, measurement error and the identified record inconsistencies; it is not a complete uncertainty bound for all lapsed trials.
- The pricing significance result is a conventional unweighted two-proportion test. Lack of statistical significance is not evidence that the segments have identical pricing sensitivity.
- The supplied 109 open texts are only 23.7% of completes. Their selection or availability is not explained. Theme counts describe this supplied set, not prevalence across all 460 respondents or all lapsed trials.
- Theme coding allows overlap. Setup counts include explicit time-consuming setup as well as lack of available time; the small-business count follows the supplied labels despite the headcount inconsistencies.
- Reconcile respondent IDs, size labels, Q1 and Q2 before relying on segment-level estimates. Different interpretations of company versus claiming-team size could explain some discrepancies, but this has not been established.
- Q3 is prompted and multi-select. A selected integration is a stated requirement, not necessarily the sole blocker or a purchase commitment. Assisted setup and trial length were not Q3 options; 'None of these' cannot be read as needing no changes.
- Q5 asks a hypothetical question, and 293/460 respondents already use supported Xero or QuickBooks. Its overall 37.6% likely-to-pay result therefore does not measure the uplift from a new integration.
- There is no Sage-specific Q5 table, no respondent-level overlap table for integration, price and workflow requirements, and no information on expense-user counts. These gaps prevent a credible integration revenue forecast.
- FreeAgent and Other software columns have small bases of 33 and 17. Comparisons involving them are especially uncertain.
- Mentions of a fortnight or trial week in the texts do not establish respondents' assigned trial duration. Verify actual trial exposure before drawing duration conclusions.
Next steps
- researchAudit respondent linkage, company size and actual trial duration; reproduce all tables and weights from respondent-level data.
- decisionKeep the standard trial at 30 days for now and make assisted setup the primary next-quarter conversion experiment.
- experimentRandomly offer guided setup to eligible new Xero and QuickBooks trials, keeping trial duration and price constant.
- researchValidate exact Sage-product demand, integration workflows and real willingness to pay with qualified prospects; obtain engineering estimates and bounded pilot commitments.
- experimentTest an alternative charging structure against current pricing, using actual offers rather than another hypothetical willingness question.
- researchObserve setup sessions and reproduce the mobile and approval failures reported in the texts.
- experimentInstrument activation and reminder response; consider a controlled 14-day versus 30-day test only after clarifying setup behaviour.
- monitorCompare respondent profiles with nonrespondents and add successful-trial evidence to the conversion analysis.
Open-text themes it coded
Insufficient time to complete setup or setup taking too long 24Missing Sage integration 15Missing FreeAgent integration 3Other or unspecified missing accounting integration 4Price too high or per-user charging unsuitable 23Approval process does not fit existing working practices 15Receipt scanning, capture or mobile upload failures 10Selected another tool 8Insufficient need or expense volume to justify an expenses app 12
The survey it planned
3 screening questions and 10 questions, as the model wrote them.
-
S1
Were you an administrator of your team's most recent Tallyroom trial, and were you involved in deciding whether to continue using it?
One answer- Yes, I was an administrator and was involved in the decision
- I was an administrator, but was not involved in the decision
- I was involved in the decision, but was not an administrator
- Neither
- Don't know
Continues if Yes, I was an administrator and was involved in the decision; I was an administrator, but was not involved in the decision
-
S2
Is the business that took this Tallyroom trial based in the UK?
One answer- Yes
- No
- Don't know
Continues if Yes
-
S3
Which statement describes your team's most recent Tallyroom trial?
One answer- It ended within the last 90 days, and we have not started a paid Tallyroom subscription
- It ended within the last 90 days, and we have started a paid Tallyroom subscription
- It ended more than 90 days ago
- It has not ended yet
- Don't know
Continues if It ended within the last 90 days, and we have not started a paid Tallyroom subscription
-
Q1
Approximately how many people does your business employ in total, including you? Please count employees across the business, not just people who would use Tallyroom.
One answer- 1–9
- 10–49
- 50–249
- 250 or more
- Don't know
-
Q2
How far did your team get during its most recent Tallyroom trial? Select the furthest stage reached.
One answer- We did not start setting it up
- We started setting it up, but did not submit an expense
- We submitted an expense, but did not approve one
- We approved an expense, but did not send one to our accounting software
- We sent an approved expense to our accounting software
- Don't know
-
Q3
Which, if any, of the following contributed to your team not starting a paid Tallyroom subscription after the trial? Select all that apply. Select 'None of these' or 'Don't know' only on its own.
Any that apply- We did not have enough need for an expenses app
- Our existing process or another product suited us better
- Tallyroom did not connect to the accounting software we needed
- Other features or capabilities we needed were missing
- Setting up or connecting Tallyroom was difficult
- Getting staff or managers to use it was difficult
- The £6 per user per month price was too high
- We could not get budget or approval to buy it
- We had concerns about reliability, security or data handling
- We were too busy to evaluate it
- We needed more time to evaluate it
- We lost track of the trial or forgot to return to it
- Another reason
- None of these
- Don't know
-
Q4
Which was the single most important reason your team did not start a paid Tallyroom subscription? If you selected reasons on the previous page, choose the most important of those.
One answer- We did not have enough need for an expenses app
- Our existing process or another product suited us better
- Tallyroom did not connect to the accounting software we needed
- Other features or capabilities we needed were missing
- Setting up or connecting Tallyroom was difficult
- Getting staff or managers to use it was difficult
- The £6 per user per month price was too high
- We could not get budget or approval to buy it
- We had concerns about reliability, security or data handling
- We were too busy to evaluate it
- We needed more time to evaluate it
- We lost track of the trial or forgot to return to it
- Another reason
- None of these
- Don't know
-
Q5
Which accounting software would Tallyroom need to send approved expenses to for your business? Select all that apply. Select 'We would not need an accounting connection' or 'Don't know' only on its own.
Any that apply- Xero
- QuickBooks
- Sage
- FreeAgent
- Zoho Books
- Microsoft Dynamics
- Other accounting software
- We would not need an accounting connection
- Don't know
-
Q6
Was there a particular missing accounting connection or feature that your team needed before it could pay for Tallyroom? If so, briefly name it and explain what you needed to do. If you selected 'Other accounting software', please name it here. You can write 'None' or 'Don't know'.
Open text -
Q7
Thinking about the 30-day trial, how did the time available compare with the time your team needed to decide whether to pay?
One answer- Much less time than we needed
- A little less time than we needed
- About the right amount of time
- A little more time than we needed
- Much more time than we needed
- We did not evaluate Tallyroom
- Don't know
-
Q8
If your team had been offered a 14-day trial instead of a 30-day trial, with the same product and £6 per user per month price, how would that have affected its chance of starting a paid subscription?
One answer- Much less likely to pay
- A little less likely to pay
- No difference
- A little more likely to pay
- Much more likely to pay
- Don't know
-
Q9
Which statement best describes your team's experience of setting up Tallyroom during the trial?
One answer- We did not attempt setup
- We completed setup without needing help
- We needed help and received enough to complete setup
- We needed help, but did not receive enough to complete setup
- We needed help, but did not ask for it
- We were unable to complete setup for a reason that help would not have resolved
- Don't know
-
Q10
Looking back, which ONE change, if any, would have made the biggest difference to your team's decision about paying for Tallyroom? This is a research question, not an offer.
One answer- A connection to the accounting software we needed
- Hands-on help from Tallyroom to set up our account and accounting connection
- A lower-priced plan, depending on its price and what it included
- A 14-day trial instead of a 30-day trial
- More than 30 days to try it
- Another product feature or improvement
- Another change
- None of these would have changed our decision
- Don't know
Sample plan
Invite one admin per UK team whose trial ended without paying in the last 90 days, using the product database. Exclude teams that have subsequently paid, suppress duplicate invitations, and recheck eligibility at fieldwork. Send a neutral invitation asking about the trial experience, describing a survey of about five minutes and a £10 voucher for completing it; do not mention sales' preferred outcome. Send up to two reminders to nonrespondents. Pilot the questionnaire on mobile and desktop to verify completion time before full launch. Randomise substantive options in Q3, retain that order in Q4, and randomise change options in Q10; keep residual options last and ordered scales fixed. Make Q6 optional. Target approximately 450 eligible completes, aiming for 150 in each of the 1–9, 10–49 and 50–249 employee bands. Use existing database size information for targeted reminders where available and Q1 to confirm quotas; do not infer employee count from Tallyroom seat count. Record and report any 250+ or unknown-size respondents separately; they do not satisfy the three band quotas. The requested precision is not generally achievable with 450 completes: under simple random sampling, 150 per band gives a worst-case 95% margin of approximately ±8 percentage points, whereas ±6 requires about 267 per band, or 801 total, before allowing for weighting. Finite-population corrections may reduce these requirements if actual eligible band populations are small and the inference is restricted to this 90-day cohort, but must be calculated from verified frame counts rather than assumed. The context suggests roughly 1,558 lapsed teams per quarter, so reaching 801 completes may be difficult. Agree either a larger sample and incentive budget, subject to frame feasibility, or wider band-level uncertainty before launch. Census invitations, quotas and nominal confidence intervals do not remove nonresponse bias.