Benchmarks / SurveyBench / What Ledgerline's accountancy practices want built next / GPT-6.1 Sol

Measured by Spring Prompt

GPT-6.1 Sol: What Ledgerline's accountancy practices want built next

The decision: which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

Verdict
✓ Sound research
Research score
100 of 100
Analysis rating
1,650
Head to head, this task
won 8 of 9

Scorecard

Numbers 11 of 11 right

QuestionIts answerTrue
What percentage of all respondents ranked AI categorisation of bank transactions first?✓37.8%37.8%
Among respondents on a paid plan, what percentage ranked more reliable bank feeds first?✓48.3%48.3%
What is the mean rank of more reliable bank feeds across all respondents (1 = highest priority, 4 = lowest), to two decimal places?✓1.991.99
What percentage of all respondents ranked AI categorisation of bank transactions last (4th)?✓42.2%42.2%
What percentage of all respondents said yes, their practice would pay £8 per user per month for AI categorisation?✓27.7%27.7%
What is the 95% margin of error, in percentage points, on the unweighted share of all respondents who ranked AI categorisation first?✓4.06 points4.06 points
Is the difference in the share ranking AI categorisation first between practices of 2-9 people and practices of 10-19 people statistically significant at the 95% level? (1 for yes, 0 for no)✓NoNo
How many of the open-text answers mention bank feeds breaking, disconnecting or arriving late?✓6774
How many of the open-text answers from respondents on a paid plan mention chasing clients for documents?✓1010
Ledgerline has 1,230 paid practices. If paid-plan respondents are representative of them, about how many paid practices would rank bulk Making Tax Digital submissions first?✓325324
What percentage of the 8,200 practice accounts invited completed the survey?✓6.68%6.7%

Traps in the data 7 of 7 handled

  • ✓AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • ✓Free-plan users are 46% of respondents but bring in no revenue; paid practices (15% of accounts, all of the revenue) put bank feeds first (48%) and bulk MTD (26%) ahead of AI (10%). The AI first-choice lead comes from free users (70% of whom rank it first). Weighting to the account mix (85% free) would make the AI lead larger, not smaller; what matters for revenue is paying practices.
  • ✓28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • ✓Every one of the 12 practices with 20 or more people ranked the client portal first.
  • ✓69% of completes came from an in-app banner seen only by users who logged in; 52% of banner respondents use Ledgerline on 5 or more days a week against 19% of email respondents and 18% of all accounts. The sample over-represents heavy users and the response rate is 6.7%.
  • ✓Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.
  • ✓The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

The analysis

Recommendation

Select more reliable bank feeds for the next half-year, explicitly on the objective of protecting paid-customer value and strengthening the core data service, rather than claiming it wins an all-account popularity vote. It has the broadest support across the observed rankings and is the leading priority among respondents from revenue-generating accounts. Begin with a short technical and customer diagnostic to identify preventable disconnections, delayed or missing transactions, and avoidable reconnection effort; distinguish these from unavoidable bank authorisation requirements. Proceed with delivery against measured reliability and customer-effort outcomes. In parallel, run limited AI purchase and portal sales-discovery tests without committing either to a full build. Keep bulk MTD as the strongest alternative for paid-practice efficiency. This is a provisional strategic choice, not a survey-proven financial optimum.

What would change it

Change the recommendation if technical diagnosis shows that the important feed problems cannot be materially improved within the half-year, or if comparative evidence shows another option has greater achievable value after development and operating costs. Examples include a real £8 AI purchase test demonstrating durable adoption, sufficient paid seats, acceptable accuracy and positive unit economics; verified larger-firm opportunities showing that a portal is a decisive buying requirement with credible incremental revenue; or measured MTD workflow savings and adoption that exceed the achievable feed benefit. If leadership instead prioritises free-account growth or conversion over existing paid-customer value, the strong free-plan AI preference warrants reconsidering the choice.

Findings

  1. Build more reliable bank feeds next if the business objective is protecting existing paid revenue and improving the core service. This is not the same as maximising first-choice preference across all accounts.

    Medium confidenceBase 548

    Among 296 paid-plan respondents, 48.3% ranked feeds first, compared with 26.4% for bulk MTD, 15.2% for the portal and 10.1% for AI. Feeds had the best paid-plan mean rank, 1.76, and only 5.1% ranked them last. Across all 548 respondents, feeds were ranked in the top two by 73.8%, had the best mean rank at 1.99, and were ranked last by only 7.5%. Paid practices generate all current revenue.

  2. CTO: do not tell the board that the survey unambiguously backs building AI next. It identifies a strong but sharply segmented demand for AI, not a consensus.

    High confidenceBase 548

    AI received 37.8% first-place votes versus 34.9% for feeds, a 2.9-point lead. But 42.2% ranked AI last, and its mean rank was 2.57 versus 1.99 for feeds. Among free-plan respondents, 70.2% ranked AI first; among paid respondents, only 10.1% did and 65.5% ranked it last. The nominal 95% margin of error around AI's unweighted first-place share is 4.06 percentage points; this does not account for recruitment bias.

  3. The sample cannot be treated as an all-account vote. Correcting only the plan imbalance would make AI substantially more popular, so the roadmap recommendation must state its commercial objective explicitly.

    Medium confidenceBase 548

    Paid accounts are 54.0% of respondents but 15% of the population; free accounts are 46.0% of respondents but 85% of the population. Applying only the population plan shares gives illustrative first-choice shares of 61.2% for AI and 23.4% for feeds. The corresponding mean ranks are approximately 1.91 for AI and 2.18 for feeds. These are sensitivity calculations, not validated population estimates, because engagement and other response biases remain.

  4. Finance director: the survey cannot provide a defensible AI add-on revenue forecast. Use stated willingness as a recruitment signal for a real purchase test, not as conversion.

    Low confidenceBase 548

    152 of 548 respondents said yes, approximately 27.7%; this comprises 102 of 252 free-plan respondents and 50 of 296 paid-plan respondents. Paid-plan willingness was 16.9%. If that rate represented all 1,230 paid practices, it would imply about 208 interested practices. Assuming every one actually purchased exactly one add-on seat gives approximately £1,663 monthly or £19,956 annual gross revenue. That is an illustrative scenario, not a forecast. The survey did not measure purchasing commitments or add-on seat counts.

  5. Head of sales: investigate the portal as a larger-firm sales opportunity, but do not prioritise the whole roadmap on this evidence alone.

    Low confidenceBase 12

    All 12 respondents from practices with 20 or more people ranked the portal first. However, only 6.6% of the 106 respondents from practices with 10–19 people did so, versus 49.1% for feeds and 30.2% for bulk MTD. Four of the largest-practice respondents spontaneously mentioned document chasing. These are existing customers, not prospective firms or lost deals.

  6. Bulk MTD is a credible next candidate for paid-practice workflow efficiency, but feeds have the stronger paid-customer priority signal.

    Medium confidenceBase 296

    Among 296 paid respondents, 26.4% ranked bulk MTD first, its mean rank was 2.23, and only 7.8% ranked it last. If representative of 1,230 paid practices, approximately 325 would rank it first. Twenty-five open answers mentioned individual MTD submission, including 22 from paid respondents, often describing concentrated deadline-period workloads.

  7. Do not infer a meaningful AI preference difference between the 2–9 and 10–19 staff groups.

    Medium confidenceBase 278

    AI was ranked first by 20.3% of 172 respondents in practices with 2–9 people and 14.2% of 106 respondents in practices with 10–19 people. Using the implied counts of 35 and 15, a two-sided pooled two-proportion test gives approximately p = 0.20, so the difference is not significant at the 95% level. This test does not establish representativeness or separate size from plan effects.

  8. Treat feed reliability and manual categorisation as related but distinct problems. AI will not repair missing or delayed source data.

    Medium confidenceBase 186

    Of 186 open answers, 67 mentioned broken, disconnected, expired or late feeds; another seven mentioned missing feed transactions without explicitly describing breakage or delay. Fifty-six mentioned manual categorisation, sometimes alongside feed failures. L511 also said existing categorisation suggestions were rarely right. The qualitative evidence supports investigating data integrity, reconnection and coding effort separately.

  9. Monitor application performance alongside the selected improvement; the four proposed features do not cover every important source of friction.

    Medium confidenceBase 186

    Twenty-six of the 186 open answers described slow loading, lag, freezing or timeouts. Twenty-five reported no slowdown. Performance complaints included client records, search, reports and deadline views, rather than only bank-feed workflows.

Caveats

  • This is a voluntary-response survey, not a probability sample. Only 548 of 8,200 invited practice accounts completed it, a 6.68% completion rate. The 83 starts that did not complete may include screened-out respondents as well as abandonment.
  • Paid practices are substantially overrepresented: 54.0% of completes versus 15% of accounts. Account-record agreement validates plan classification, not sample representativeness.
  • Frequent users are also overrepresented: 41.2% reported use on at least five days a week versus 18% in population records; only 12.0% reported use less than weekly versus 35% in population records. The survey and records refer to different periods and use self-report versus recorded activity.
  • Recruitment channel is associated with engagement: 69% of completes came from the in-app banner; 52% of banner respondents versus 19% of email respondents reported use on at least five days a week. Channel effects cannot be separated from engagement or plan effects using the supplied tables.
  • The 4.06-point margin of error is the conventional unweighted binomial calculation, 1.96 times the square root of p(1-p)/548. It is not a valid bound on total population-estimation error for this self-selected sample.
  • Plan-only weighting is an illustrative sensitivity check. Respondent-level joint plan and usage data, harmonised population measures and assessment of nonresponse are needed for a stronger weighting analysis. Weighting cannot remove unobserved selection bias.
  • Practice size, plan and engagement may be associated. The supplied marginal tables do not support attributing differences independently to practice size.
  • Rankings force priorities among four supplied options. They do not measure absolute benefit, actual adoption, time saved, retention effects, delivery cost or preferences for omitted improvements. Randomising option order helps with order effects but does not resolve these limitations.
  • Q3 measures hypothetical willingness at one price, not purchases, price elasticity or billable seats. Staff count is not Ledgerline user count. Free-plan eligibility for the add-on, possible plan conversion and net revenue after delivery costs are unspecified.
  • The portal result for practices with 20 or more staff rests on 12 existing customers. It does not establish demand among prospective larger firms or demonstrate that a portal would win deals.
  • Open-text themes are researcher-coded spontaneous mentions from 186 respondents, not prevalence estimates for all 548 respondents or all accounts. Multiple themes can occur in one answer. Manual-coding complaints do not by themselves establish that AI is the preferred solution.
  • The bankfeed_mentions answer uses the narrow requested definition: 67 answers explicitly describe breakage, disconnection, expiry/re-authorisation or delay. Seven additional answers describe missing transactions only; the broader feed-reliability count is 74. The documents_paid answer counts document or record chasing, not requests to reauthorise bank access.
  • Some feed reauthorisation may be required by banks or regulation rather than caused by a Ledgerline defect. Respondents' accounts establish friction, not technical responsibility or the feasibility of eliminating it.
  • Fielding ran from 7 to 20 September 2026. MTD and document-chasing priorities may vary around tax and VAT deadlines; these results are a snapshot.

Next steps

  • decisionAgree and document the roadmap objective and investment criteria before presenting the decision to the board.
  • researchAudit feed failures and observe affected practices completing reconciliation and reconnection workflows.
  • experimentDeliver and evaluate the feed-reliability improvement in a staged rollout.
  • experimentRun an AI categorisation prototype with a genuine £8-per-user purchasing decision across free and paid practices.
  • researchInvestigate the portal in larger-firm sales opportunities, including lost deals and prospective customers.
  • researchMeasure bulk MTD workload through a deadline period and test a batch-submission workflow.
  • researchAnalyse respondent-level data and recruit underrepresented infrequent users for targeted follow-up.
  • monitorTrack core application performance and paid-customer outcomes alongside the feed work.

Open-text themes it coded

Bank feeds breaking, disconnecting, expiring, requiring reconnection or arriving late 67Missing bank-feed transactions without an explicit mention of breakage, disconnection or delay 7Manual bank-transaction categorisation and inadequate coding automation 56Waiting for or chasing clients' documents and records 15Individual MTD submissions and lack of bulk filing 25Slow application performance, lag, freezing or timeouts 26No reported slowdown or problem 25

The survey it planned

2 screening questions and 8 questions, as the model wrote them.

  1. S1

    Are you answering on behalf of a UK accountancy or bookkeeping practice that currently uses Ledgerline?

    One answer
    • Yes
    • No
    • Not sure

    Continues if Yes

  2. S2

    Which best describes your involvement in choosing software for your practice?

    One answer
    • I make the final decision
    • I share responsibility for the decision
    • I advise or make recommendations to the decision-makers
    • I use the software but do not influence the decision
    • I am not involved

    Continues if I make the final decision; I share responsibility for the decision; I advise or make recommendations to the decision-makers

  3. Q1

    Which Ledgerline plan does your practice currently use?

    One answer
    • Free plan
    • Paid plan
    • Not sure
  4. Q2

    How many people currently work in your practice, including owners, employees and contractors who regularly work for the practice? Count people, not full-time equivalents.

    One answer
    • 1 person — sole practitioner
    • 2–9 people
    • 10 or more people
    • Not sure
  5. Q3

    Over the past four weeks, how often has anyone in your practice used Ledgerline to do work, rather than just logging in? Please answer for the practice as a whole.

    One answer
    • Every or almost every working day
    • Several days a week
    • About once a week
    • Less than once a week
    • Not at all
    • Not sure
  6. Q4

    Over the past four weeks, which of these, if any, have slowed your practice down when working with Ledgerline? Select up to three. Select 'Nothing slowed us down', 'We have not used Ledgerline in the past four weeks' or 'Not sure' on its own.

    Any that apply
    • Categorising or checking bank transactions
    • Bank-feed connections breaking or needing to be reconnected
    • Requesting, chasing or collecting client documents
    • Preparing or sending Making Tax Digital submissions one client at a time
    • Keeping client records up to date
    • Managing or checking deadlines
    • Finding information or understanding how to complete a task
    • Slow loading, errors or other technical problems apart from bank-feed connections
    • Something else
    • Nothing slowed us down
    • We have not used Ledgerline in the past four weeks
    • Not sure
  7. Q5

    If something slowed your practice down, briefly describe the issue that had the biggest impact and what it prevented or delayed. A recent example would help. Please do not include client names or confidential information. You can leave this blank.

    Open text
  8. Q6

    Ledgerline can build one of the following improvements in the next six months. Rank all four in the order your practice would want them built: 1 = highest priority and 4 = lowest priority. Put improvements your practice would not use below those it would use. There are no right or wrong answers.

    Ranking
    • AI categorisation of bank transactions — AI suggests transaction categories for your practice to review and approve.
    • More reliable bank feeds — reduce broken connections and the need to reconnect bank feeds.
    • Client document portal — request documents from clients and let them upload those documents securely.
    • Bulk Making Tax Digital submissions — review and send eligible MTD submissions for multiple clients together, rather than one client at a time.
  9. Q7

    Suppose AI categorisation were offered as an optional paid add-on to your practice's current Ledgerline plan. It would suggest bank transaction categories for your practice to review and approve. The add-on would be charged for each Ledgerline user in your practice each month, in addition to any existing plan charge. Would your practice consider paying for this add-on?

    One answer
    • Definitely would
    • Probably would
    • Might or might not
    • Probably would not
    • Definitely would not
    • Not sure
  10. Q8

    What is the highest additional amount your practice would be willing to pay for AI categorisation, in pounds per Ledgerline user per month, excluding VAT? Remember that this amount would be charged for every Ledgerline user in your practice, on top of any existing plan charge. Enter 0 if your practice would not pay anything. You can leave this blank if you cannot estimate an amount.

    Number

    Routing Shown only if Q7 is: Definitely would; Probably would; Might or might not; Probably would not

Sample plan

Invite the admin of all 8,200 practice accounts by email and show an in-app invitation to users when they log in. State that the survey takes about five minutes, concerns Ledgerline's next improvement, and accepts one response per practice; do not mention the CTO's preference. Use account-linked invitation tokens and authenticated in-app links that resolve to the same practice ID. Keep the first eligible completed response per practice and block subsequent submissions; maintain the same rule across both channels. Link responses to account plan and, where permitted, recent account activity; disclose that linkage in the survey introduction. Use account billing data as the authoritative plan classification, retaining Q1 to identify discrepancies. Target 550 eligible, unique-practice completes. Aim for 275 paid and 275 free practices, deliberately oversampling paid practices because they provide all current revenue. Across that total, aim for 184 sole practitioners, 183 practices with 2–9 people and 183 with 10 or more people. Monitor the plan-by-size cells, close filled quotas and send targeted reminders to underrepresented groups. Check feasibility against available account data; do not silently replace an unachievable quota or count unknown-size responses towards a known-size target. The requested precision is not generally achievable with 550 completes: approximately 385 independently sampled completes per size group, or 1,155 total, are needed for a worst-case proportion at ±5 percentage points and 95% confidence before finite-population correction. At about 183 completes per group, the corresponding unweighted margin is about ±7.2 points before that correction. Obtain agreement before launch either to accept lower subgroup precision or to increase the sample. If reliable population counts by size are available, recalculate finite-population requirements separately for each group. Weighting can reduce effective sample size, and voluntary banner/email participation introduces selection bias; nominal confidence intervals do not guarantee population accuracy. Pilot on mobile and desktop with eligible respondents to verify comprehension and completion within five minutes. Randomise Q4's substantive options while keeping the final four fixed; randomise Q6's starting order independently. Enforce Q4's three-choice limit and exclusive responses, and accept non-negative numeric amounts in Q8. Show no AI-payment questions before Q6.