Benchmarks / SurveyBench / What Ledgerline's accountancy practices want built next / Claude Opus 5.5

Measured by Spring Prompt

Claude Opus 5.5: What Ledgerline's accountancy practices want built next

The decision: which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

Verdict
✓ Sound research
Research score
100 of 100
Analysis rating
1,646
Head to head, this task
won 10 of 11

Scorecard

Numbers 11 of 11 right

QuestionIts answerTrue
What percentage of all respondents ranked AI categorisation of bank transactions first?✓37.8%37.8%
Among respondents on a paid plan, what percentage ranked more reliable bank feeds first?✓48.3%48.3%
What is the mean rank of more reliable bank feeds across all respondents (1 = highest priority, 4 = lowest), to two decimal places?✓1.991.99
What percentage of all respondents ranked AI categorisation of bank transactions last (4th)?✓42.2%42.2%
What percentage of all respondents said yes, their practice would pay £8 per user per month for AI categorisation?✓27.7%27.7%
What is the 95% margin of error, in percentage points, on the unweighted share of all respondents who ranked AI categorisation first?✓4.1 points4.06 points
Is the difference in the share ranking AI categorisation first between practices of 2-9 people and practices of 10-19 people statistically significant at the 95% level? (1 for yes, 0 for no)✓NoNo
How many of the open-text answers mention bank feeds breaking, disconnecting or arriving late?✓7474
How many of the open-text answers from respondents on a paid plan mention chasing clients for documents?✓1010
Ledgerline has 1,230 paid practices. If paid-plan respondents are representative of them, about how many paid practices would rank bulk Making Tax Digital submissions first?✓325324
What percentage of the 8,200 practice accounts invited completed the survey?✓6.7%6.7%

Traps in the data 7 of 7 handled

  • ✓AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • ✓Free-plan users are 46% of respondents but bring in no revenue; paid practices (15% of accounts, all of the revenue) put bank feeds first (48%) and bulk MTD (26%) ahead of AI (10%). The AI first-choice lead comes from free users (70% of whom rank it first). Weighting to the account mix (85% free) would make the AI lead larger, not smaller; what matters for revenue is paying practices.
  • ✓28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • ✓Every one of the 12 practices with 20 or more people ranked the client portal first.
  • ✓69% of completes came from an in-app banner seen only by users who logged in; 52% of banner respondents use Ledgerline on 5 or more days a week against 19% of email respondents and 18% of all accounts. The sample over-represents heavy users and the response rate is 6.7%.
  • ✓Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.
  • ✓The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

The analysis

Recommendation

Build more reliable bank feeds next half-year. Define the scope as cutting failures that Ledgerline can fix, plus proactive alerts when a feed drops, a smoother client re-authorisation flow, and duplicate and gap handling. This is the clear first choice of paying practices (48.3% first, mean rank 1.76, n=296), the least-deprioritised option overall (7.5% last), and the most common complaint in the open text (74 of 186). It also removes a large share of the manual coding burden that drives demand for AI. Before committing the full scope, confirm from connection telemetry how much of the failure rate is within Ledgerline's control. Treat bulk MTD as the leading candidate for the following half-year, since it is second among paid practices and concentrated in multi-person firms.

Answers to the stakeholder questions:

(1) CTO: No. AI's first-place lead (37.8% vs 34.9%) is within the margin of error. It is the option most often ranked last (42.2%), and its support comes from free single-user practices that bring in no revenue. Only 10.1% of paid practices rank it first and 65.5% rank it last. The board should hear that the survey backs bank feeds.

(2) Finance director: The survey cannot give a credible add-on revenue figure. It does not record users per practice, it is not known whether free accounts could buy the add-on, and stated intent overstates purchase. Only 16.9% of paid respondents said yes. Taken at face value, the 40.5% yes among free respondents would imply a ceiling of roughly £23k a month from one-user free accounts (about 2,800 × £8). That is a ceiling, not a forecast. A priced fake-door test is needed before any revenue figure is used.

(3) Head of sales: Not on this evidence. The '100% of the biggest practices' result is 12 respondents, contradicted by 10-19 person practices (6.6% portal first). Run targeted research with large customers and lost or prospective large firms before prioritising a portal.

What would change it

Bank feeds would lose first place if connection telemetry showed most failures are outside Ledgerline's control, such as regulatory consent expiry or bank-side outages, so that a half-year of work could not materially cut breakages. In that case, ship the cheaper alerting and re-authorisation improvements and give the main build to bulk MTD, which is second among paid practices. AI categorisation would move up if the business explicitly chose free-to-paid conversion as the half-year goal, and a priced fake-door or pre-order test showed meaningful real conversion or add-on purchase. The client portal would move up if research with at least 15–20 large customers and lost prospects showed document requests are a decisive buying or churn factor. A weighted re-analysis by usage frequency that reversed the paid-practice ranking would also change it, though that is unlikely given the size of the gap.

Findings

  1. Paid practices, which bring in all of Ledgerline's revenue, put more reliable bank feeds clearly first. This is the strongest single signal for the build decision.

    High confidenceBase 296

    Paid plan (base 296): bank feeds ranked first by 48.3% (±5.7pp) and last by only 5.1%, mean rank 1.76. Next best was bulk MTD at 26.4% first and mean 2.23; the 21.9pp gap is far outside the margin of error. AI categorisation was ranked first by 10.1% of paid respondents and last by 65.5%. Paid response rate was about 24% (296 of 1,230), so this group is reasonably well covered.

  2. Bank feeds is the only option that almost nobody deprioritises. It has the best mean rank across the whole sample.

    High confidenceBase 548

    All respondents (base 548): mean rank 1.99, against 2.57 for AI, 2.67 for MTD and 2.77 for client portal. Bank feeds was ranked first or second by 73.8% and last by 7.5%. No size group ranked it last more than 9.3% of the time.

  3. AI categorisation's first-place lead is not a real lead. It comes almost entirely from free single-user practices, and the option is the most polarising of the four.

    High confidenceBase 548

    All respondents: AI first 37.8% (±4.1pp) against bank feeds 34.9%. The 2.9pp gap is inside the ±7.1pp margin for comparing two shares from the same sample, so it is not significant. 42.2% ranked AI last, the most of any option. Free plan (base 252): 70.2% ranked AI first, mean 1.66. Just-me practices (base 258): 60.9% first. Paid plan: 10.1% first, 65.5% last. The unweighted total is a sample mix (54% paid) that matches neither the revenue base (100% paid) nor the account base (15% paid).

  4. Broken, expiring and late bank feeds are the most common complaint in the open text, especially from paid practices. They also cause much of the manual coding that AI categorisation would target.

    Medium confidenceBase 186

    74 of 186 open-text answers (40%) mention feeds breaking, disconnecting, expiring or arriving late; 51 of these are from paid respondents. Of the 56 answers about manual coding, 25 also blame feed outages for the backlog (e.g. L126, L235, L378, L472). Only 16 coding mentions come from paid respondents, mostly tied to feed outages. Respondents also ask for warnings when feeds drop (L018, L051, L417) and for an easier re-authorisation process (L028, L105, L342).

  5. Bulk MTD submission is a solid second priority for paid, multi-person practices. It is a candidate for the following half-year rather than a competitor for this one.

    Medium confidenceBase 296

    Paid: 26.4% ranked it first (±5.0pp), about 325 of 1,230 paid practices (range roughly 263–386) if respondents are representative. Ranked first by 30.2% of 10-19 person practices (base 106) and 22.7% of 2-9 (base 172); last by only 7.8% of paid. 25 open-text answers raise it, 22 of them from paid respondents, focused on quarter-end and the 7th-of-month peak.

  6. Stated willingness to pay £8/user/month for AI categorisation is weakest among the practices that already pay.

    Medium confidenceBase 548

    Paid plan (base 296): 16.9% yes, 57.4% no, 25.7% not sure. Free plan (base 252): 40.5% yes, but these are one-user accounts that currently pay nothing, and it is not established whether the free plan could buy add-ons. Yes falls as practice size rises: 36.8% (just me) to 16.7% (20+, base 12). Stated intent usually overstates real purchase.

  7. The survey cannot show that large firms want a client portal. Twelve respondents is far too few to set strategy for winning larger firms.

    Low confidenceBase 12

    20+ people (base 12): 12 of 12 ranked client portal first. The margin of error on 12 respondents is very wide, and the next size band, 10-19 people (base 106), ranked the portal first only 6.6% of the time. Document-chasing appears in 15 of 186 open-text answers, including all 4 from 20+ firms. The survey covers existing customers only, not the larger prospects sales wants to win.

  8. Application slowness is a substantial unprompted complaint that none of the four options addresses.

    Medium confidenceBase 186

    26 of 186 open-text answers describe slow page loads, search, reports or dashboards, e.g. L279 (slower since growing to 300 clients), L093 and L120 (slow in the morning), L041 (slow at peak). A further 25 answers said nothing slows them down.

  9. Practice-size differences in AI support among multi-person practices are not meaningful. The real divide is free versus paid, and sole practitioners versus multi-person practices.

    High confidenceBase 278

    AI ranked first by 20.3% of 2-9 person practices (base 172) and 14.2% of 10-19 person practices (base 106): z≈1.3, not significant at 95%. The free/paid gap (70.2% vs 10.1%) is very large.

Caveats

  • The sample does not mirror the customer base. Paid practices are 54% of respondents but 15% of accounts (about 24% of paid practices responded, against about 3.6% of free). Totals are therefore a mix that represents neither revenue nor accounts. Weighted to the account mix, AI would be about 61% first, but that weight comes almost entirely from non-paying free accounts.
  • Heavy users are over-represented: 41.2% use Ledgerline 5+ days a week against 18% of accounts. The in-app banner drove this (52% of banner respondents vs 19% of email respondents). Views of light and lapsed users, who make up 35% of accounts, are thin.
  • The overall response rate is 6.7% (548 of 8,200). Respondents chose to take part, so practices with strong feelings may be over-represented. Figures are unweighted.
  • Willingness to pay is stated intent to a hypothetical question. It typically overstates real purchase several-fold and cannot be turned into revenue directly.
  • The survey did not ask how many Ledgerline users each practice has, so per-user revenue cannot be calculated. Practice headcount is not the same as paid seats.
  • The 20+ person segment has only 12 respondents. Its 100% client-portal result cannot support decisions, and larger prospects who are not yet customers were not surveyed at all.
  • The ranking forced a choice among four pre-set options. Slowness, raised unprompted by 26 open-text respondents, was not an option.
  • Some feed failures (90-day consent expiry, bank security changes, individual bank APIs) may be partly outside Ledgerline's control. How much reliability engineering can actually improve is not yet known.
  • Open-text coding was done by a single analyst. Theme counts are indicative only, and only 186 of 548 respondents answered Q1.

Next steps

  • decisionCommit 'more reliable bank feeds' as the next half-year build. The scope covers failure reduction, proactive dropped-feed alerts, a guided client re-authorisation flow, and duplicate and gap handling. Tell the board it rests on paid-practice rankings, not the unweighted first-place count.
  • researchBefore final scoping, analyse connection telemetry by bank and failure cause (consent expiry, bank API errors, Ledgerline-side errors) and link it to support tickets and paid churn.
  • experimentRun a priced fake-door or pre-order test of AI categorisation at £8/user/month, shown separately to free and paid practices, with clear eligibility rules for free accounts.
  • researchInterview about 15–20 practices with 20+ staff, plus recently lost or prospective large firms, about document collection and their buying criteria. Review sales loss reasons as well.
  • researchSize bulk MTD demand from submission logs: the number of paid practices filing many VAT returns in the same deadline window.
  • monitorTrack page-load, search and report latency at the 95th percentile, especially at peak times such as mornings, month-end and January, and for practices with many clients.
  • researchIn future surveys, use quotas or weighting by plan and usage frequency, ask for the number of Ledgerline users, and report paid-practice results as the headline.

Open-text themes it coded

Bank feeds breaking, disconnecting, expiring or arriving late 74Manual coding/categorising of bank transactions 56MTD/VAT returns filed one client at a time 25Chasing clients for documents and records 15Slow application performance (page loads, search, reports) 26Nothing slows them down 25

The survey it planned

3 screening questions and 8 questions, as the model wrote them.

  1. S1

    Which of these best describes the organisation you work for?

    One answer
    • An accountancy practice based in the UK
    • A bookkeeping practice based in the UK
    • An accountancy or bookkeeping practice based outside the UK
    • The in-house finance team of a business or charity
    • Other

    Continues if An accountancy practice based in the UK; A bookkeeping practice based in the UK

  2. S2

    Does your practice currently use Ledgerline?

    One answer
    • Yes, we use it now
    • We used it in the past but have stopped
    • No
    • Not sure

    Continues if Yes, we use it now

  3. S3

    What is your role in choosing the software your practice uses?

    One answer
    • I make the final decision
    • I share the decision with others
    • I recommend or influence software choices but don't make the final decision
    • I'm not involved in software choices

    Continues if I make the final decision; I share the decision with others; I recommend or influence software choices but don't make the final decision

  4. Q1

    How many people work in your practice, including you? Count everyone, part-time or full-time, not just the people who use Ledgerline.

    One answer
    • Just me
    • 2 to 4
    • 5 to 9
    • 10 to 24
    • 25 or more
  5. Q2

    Thinking about the last month, what, if anything, has slowed you or your team down most when using Ledgerline? Please describe it in your own words. (Optional)

    Open text
  6. Q3

    In the last month, which of these, if any, have slowed you or your team down when using Ledgerline? Select all that apply.

    Any that apply
    • Bank feed connections breaking or needing to be reconnected
    • Categorising bank transactions
    • Chasing clients for documents or records
    • Submitting Making Tax Digital returns one client at a time
    • Keeping track of client deadlines
    • Finding information in client records
    • Setting up new clients
    • Something else
    • None of these
  7. Q4

    Ledgerline is considering four improvements. Please rank them from 1 (most valuable to your practice) to 4 (least valuable to your practice).

    Ranking
    • More reliable bank feeds: fewer broken connections and less time spent reconnecting them
    • AI categorisation of bank transactions: Ledgerline suggests a category for each transaction, and you review and approve it
    • A client portal: clients upload the documents you request in one place
    • Bulk Making Tax Digital submissions: submit returns for many clients in one go
  8. Q5

    How much difference would the improvement you ranked first make to your practice's workload?

    Scale

    Scale 1-5: No real difference to A very large difference

  9. Q6

    Suppose Ledgerline offered AI categorisation of bank transactions as an optional add-on. Ledgerline would suggest a category for each transaction, and you would review and approve it. The add-on would be charged on top of your current plan. How likely is it that your practice would pay for it?

    One answer
    • Definitely would pay
    • Probably would pay
    • Might or might not pay
    • Probably would not pay
    • Definitely would not pay
    • Don't know
  10. Q7

    What is the most your practice would pay for the AI categorisation add-on, per user per month, before VAT?

    One answer
    • Less than £2
    • £2 to £4.99
    • £5 to £9.99
    • £10 to £14.99
    • £15 or more
    • Not sure

    Routing Shown only if Q6 is: Definitely would pay; Probably would pay; Might or might not pay

  11. Q8

    Roughly how many clients does your practice currently manage in Ledgerline?

    One answer
    • 1 to 10
    • 11 to 50
    • 51 to 150
    • 151 to 400
    • More than 400
    • Not sure
Sample plan

WHO IS INVITED: all 8,200 practice accounts. Every account admin receives an email invitation with a unique link tied to the account ID. The in-app banner is shown to all users at login, and its link also carries the logged-in account ID. The banner is not shown to anyone who has already responded for that practice. Plan (free or paid), number of seats, tenure and usage data are attached from account records through the account ID, so plan is not asked in the survey. Send one email reminder after 5 days and run fieldwork for about 2 weeks.

SURVEY DESIGN: the survey has 3 screener questions and 8 questions, and median completion is about 4 minutes. The order of options is randomised in Q3 (except 'Something else' and 'None of these', which stay at the end) and in Q4. This matters because the CTO prefers AI categorisation: the invitation, introduction and question wording must not signal any preferred option, and AI must not always appear first.

ONE RESPONSE PER PRACTICE: responses are de-duplicated by account ID. If a practice sends more than one complete, the response from the most senior decision-maker is kept, in this order: S3 'final decision', then 'shared', then 'influence'. If two are equally senior, the earliest complete is kept.

DATA QUALITY: remove speeders (anyone who finishes in under one-third of the median completion time) and respondents whose Q4 ranking matches the on-screen order and who also finish unusually fast. A formal attention-check item was left out to keep within the 8-question limit.

COMPLETES AND QUOTAS (target about 550): (a) Practice size, using Q1 grouped as Sole (Just me) / 2 to 9 / 10 or more: a minimum of about 180 completes per band. Weighted invitations and reminders are used to boost the 10-or-more band, which will be the hardest to fill. (b) Plan: a minimum of 250 paid-plan completes, out of 1,230 paid practices. That gives about ±5.5pp for paid practices overall after finite-population correction. The remaining roughly 300 completes come from free-plan practices.

PRECISION CONFLICT, which must be resolved with Product before launch: ±5pp at 95% confidence (worst case, p = 50%) needs about 385 completes per size band, so about 1,150 in total. That is roughly double the 550 budget. A small band would need fewer completes because of the finite-population correction (n = 385 / (1 + 384 / N)). For example, a band of 500 practices would need about 220, but that means about 44% of that band responding, which is unrealistic. With about 183 per band, the margin of error is about ±7.2pp. If 550 completes fall naturally (with no size quotas), the 10-or-more band will probably have only a few dozen completes and a margin of ±12pp or worse. Options for Product: (1) Keep 550 with the size quotas above and accept about ±7pp per band. (2) Raise the target to about 1,150 completes, using the full invitation base, incentives such as a prize draw and extra reminders. (3) Keep ±5pp only for the comparison that drives the decision, paid practices, and accept wider margins by size. Before launch, check the number of accounts in each size band using seat counts as a proxy. Seat counts understate size for free-plan practices, because the free plan allows only one user.

WEIGHTING: total-population results are weighted back to the account base by plan and size. A second, revenue-weighted view weights paid practices by paid seats, because paid plans bring in all of the revenue.

COVERAGE LIMITATION: the survey reaches only current users. Practices that left Ledgerline, possibly because of broken bank feeds, are not represented. Churn reasons and support-ticket data should be reviewed alongside the survey results.