Benchmarks / SurveyBench / What Ledgerline's accountancy practices want built next / Claude Haiku 5.5

Measured by Spring Prompt

Claude Haiku 5.5: What Ledgerline's accountancy practices want built next

The decision: which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

Verdict
✗ Not sound
Research score
84 of 100
Analysis rating
1,412
Head to head, this task
won 7 of 9

Why it is not sound

The survey

  • Answer bands that overlap or leave gaps

The analysis

  • Reported a material finding the data does not support
  • Missed a trap: AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • Missed a trap: 28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • Proposed a mistaken next step: Forecast AI add-on revenue from the share who said they would pay £8 per user per month

Findings the judges found unsupported

  • AI categorisation's appeal is concentrated in the free plan, which brings in no revenue. Its stated paid demand is small. On the most defensible reading, about 208 paid practices would say yes at GBP 8 per user per month, which is roughly GBP 20k a year at one seat each. This is a floor, not a forecast, because seat counts are not in the survey.The £20k annual revenue 'floor' converts hypothetical stated willingness to pay into revenue. Neither actual purchases nor a minimum number of purchasers is established. One seat per purchaser is a minimum seat assumption, not a revenue floor; actual revenue could be zero. The caveats about stated intent do not make this lower-bound claim valid.

Scorecard

Numbers 11 of 11 right

QuestionIts answerTrue
What percentage of all respondents ranked AI categorisation of bank transactions first?✓37.8%37.8%
Among respondents on a paid plan, what percentage ranked more reliable bank feeds first?✓48.3%48.3%
What is the mean rank of more reliable bank feeds across all respondents (1 = highest priority, 4 = lowest), to two decimal places?✓1.991.99
What percentage of all respondents ranked AI categorisation of bank transactions last (4th)?✓42.2%42.2%
What percentage of all respondents said yes, their practice would pay £8 per user per month for AI categorisation?✓27.7%27.7%
What is the 95% margin of error, in percentage points, on the unweighted share of all respondents who ranked AI categorisation first?✓4.1 points4.06 points
Is the difference in the share ranking AI categorisation first between practices of 2-9 people and practices of 10-19 people statistically significant at the 95% level? (1 for yes, 0 for no)✓NoNo
How many of the open-text answers mention bank feeds breaking, disconnecting or arriving late?✓7274
How many of the open-text answers from respondents on a paid plan mention chasing clients for documents?✓1010
Ledgerline has 1,230 paid practices. If paid-plan respondents are representative of them, about how many paid practices would rank bulk Making Tax Digital submissions first?✓325324
What percentage of the 8,200 practice accounts invited completed the survey?✓6.7%6.7%

Traps in the data 5 of 7 handled

  • ✗AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • ✓Free-plan users are 46% of respondents but bring in no revenue; paid practices (15% of accounts, all of the revenue) put bank feeds first (48%) and bulk MTD (26%) ahead of AI (10%). The AI first-choice lead comes from free users (70% of whom rank it first). Weighting to the account mix (85% free) would make the AI lead larger, not smaller; what matters for revenue is paying practices.
  • ✗28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • ✓Every one of the 12 practices with 20 or more people ranked the client portal first.
  • ✓69% of completes came from an in-app banner seen only by users who logged in; 52% of banner respondents use Ledgerline on 5 or more days a week against 19% of email respondents and 18% of all accounts. The sample over-represents heavy users and the response rate is 6.7%.
  • ✓Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.
  • ✓The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

The analysis

Recommendation

Build more reliable bank feeds next half-year. Paid practices, which bring in all revenue, rank it first (48.3%) and it has the best mean rank among paid respondents (1.76). It is also the most-mentioned complaint (72 of 186 open-text answers). Feed failures also drive much of the manual coding that AI categorisation would target. Do not present the survey as a clear backing for AI categorisation: its first-place lead overall is not significant against feeds and comes from free-plan users who do not pay. Do not prioritise the client portal on this evidence. Treat AI categorisation as the main alternative for the free base. Before committing, test its paid demand with a priced offer and check that free accounts can buy it. Bulk MTD is the runner-up for paid retention and should be costed. Answers to stakeholders: CTO, not proven, since the lead is not significant against feeds and depends on weighting. Finance director, a floor of about GBP 20k a year from paid practices at one seat each, with no validated central estimate. Head of sales, the survey cannot support prioritising the portal for larger firms.

What would change it

Choose AI categorisation instead if: (a) a priced test shows meaningful paid conversion for the add-on, for example at least 10% of paid accounts shown the offer buying it, and free accounts can buy it; or (b) feed telemetry shows most breaks are bank-side or consent-related and cannot be reduced by Ledgerline, so the feed work would not change outcomes. Choose bulk MTD if paid-practice churn or support data show VAT submission as the main reason practices leave, or if a re-run of Q2 with paid respondents by size confirms 10-19 person practices rank it first. Revisit the portal only if win/loss evidence shows document collection is a real reason for winning or losing larger firms.

Findings

  1. Paid practices, the 15% of accounts that bring in all revenue, rank more reliable bank feeds first. AI categorisation is their last choice. Reliable bank feeds is the best-supported build for the revenue base.

    Medium confidenceBase 296

    Among the 296 paid-plan respondents, 48.3% rank bank feeds first (95% margin about +/-5.7 points) and mean rank is 1.76. AI categorisation is first for 10.1% (+/-3.4) and last for 65.5%. Bulk MTD is first for 26.4% (+/-5.0). Bank feeds are first for 43.0% of 2-9 person practices and 49.1% of 10-19 person practices.

  2. The CTO's claim is literally true of the raw sample but is not a clear result. AI categorisation has the most first-place votes (37.8%, about 207 respondents) but is not statistically separable from bank feeds (34.9%, about 191). The lead depends on weighting, and it comes almost entirely from free-plan users.

    Medium confidenceBase 548

    The unweighted gap is 2.9 points. Its 95% margin is about +/-7 points, because the two shares come from the same respondents. AI is first for 70.2% of free-plan respondents (+/-5.6) and 10.1% of paid. The sample is 54% paid against 15% of accounts. Weighting plan mix to the account base gives AI 61.2% first and bank feeds 23.4%, so the result is sensitive to weighting. Mean rank on the same weighting is 1.91 for AI and 2.18 for feeds.

  3. AI categorisation's appeal is concentrated in the free plan, which brings in no revenue. Its stated paid demand is small. On the most defensible reading, about 208 paid practices would say yes at GBP 8 per user per month, which is roughly GBP 20k a year at one seat each. This is a floor, not a forecast, because seat counts are not in the survey.

    Low confidenceBase 296

    Q3 'yes' is 16.9% among paid respondents (50 of 296), so about 208 of 1,230 paid practices (16.9%). 'Not sure' is 25.7% for paid respondents. Free-plan 'yes' is 40.5% (102 of 252), but the add-on's availability to free accounts is not established. Stated purchase intent usually overstates actual purchase, and no haircut is validated here.

  4. The head of sales' portal question cannot be answered from this survey. Portal first-place support is 11.7% overall and 6.6% among 10-19 person practices. The only 20+ person group is 12 respondents, all of whom rank the portal first, which is too small to support a conclusion. Four 20+ respondents mention chasing documents in open text, so the pain is real but not yet linked to winning larger firms.

    Low confidenceBase 12

    Portal first: 8.1% of 1-person practices (21), 14.0% of 2-9 (24), 6.6% of 10-19 (7), 100% of 20+ (12). Portal is first for only 15.2% of paid respondents. Document-chasing mentions in open text are 15 in total (8% of 186), of which 10 are from paid respondents.

  5. Bank feed failure is the most frequent complaint in the open text. Many coding complaints are downstream of feed failures, so reliable feeds would also reduce part of the manual-coding burden.

    Medium confidenceBase 186

    72 of 186 open-text answers (38.7%) describe feeds breaking, disconnecting, expiring or arriving late. 50 of these are from paid respondents. 56 answers (30.1%) describe manual categorisation. 25 answers mention both, so 103 of 186 (55.4%) name one or both. Many coding complaints describe clearing a backlog after a feed outage.

  6. Bulk MTD submission is the strongest alternative for paid practices, especially 10-19 person firms. It is the runner-up if the board prioritises paid retention. About 325 paid practices would rank it first if the paid sample is representative.

    Medium confidenceBase 296

    26.4% of paid respondents rank bulk MTD first (+/-5.0, n=296), and 30.2% of 10-19 person practices do (n=106). 22 of the 25 open-text MTD complaints come from paid respondents. Multiplying 26.4% by 1,230 paid practices gives about 325.

  7. The size difference in AI first-place support between 2-9 and 10-19 person practices is not statistically significant. Size does not justify targeting AI categorisation differently by segment.

    Medium confidenceBase 278

    20.3% (35 of 172) against 14.2% (15 of 106). The difference is 6.1 points. A two-proportion test gives z of about 1.3 and p of about 0.20.

  8. The sample is not representative of the 8,200 accounts, so unweighted totals must not be quoted as the market view. It over-represents paid practices and heavy users and reflects only 6.7% of accounts.

    High confidenceBase 548

    Paid share is 54% of completes against 15% of accounts. Free-plan respondents are about 3.6% of their population (252 of about 6,970). Five or more days of use is 41% of respondents against 18% of accounts. The email invite produced 19% heavy users, against 52% from the in-app banner.

Caveats

  • Q2 is a forced ranking of four options. Ranks show relative priority, not how much each matters, and the ranking cannot show whether a practice would switch away from Ledgerline.
  • The sample is self-selected from in-app and email invitations. Only 548 of 8,200 accounts (6.7%) responded, with one response per practice. The reported margins of error assume random sampling, so true uncertainty is larger.
  • Tables are unweighted. Paid practices are 54% of respondents against 15% of accounts, and heavy users are 41% against 18%. Plan-weighted figures are my own calculations and assume that free and paid respondents are representative within each plan. Usage was not cross-tabulated with Q2, so usage weighting is not possible.
  • The 1,230 paid-practice estimates (325 for MTD and about GBP 20k for AI revenue) assume paid respondents represent paid practices. The survey does not record seat counts, so revenue is a floor at one seat per practice.
  • Q3 asks about intent to pay for AI categorisation at GBP 8 per user per month. Stated intent overstates real purchase, and the survey does not show whether the add-on would be sold to free-plan accounts.
  • The bank-feed count (72) uses this definition: feeds broken, disconnected, expired, needing re-authorisation, arriving late, or transactions missing from the feed. Two answers that mention missing bank transactions without naming the feed (L322, L492) are excluded. Some feed failures may be caused by banks or by client re-consent rather than by Ledgerline, which the survey cannot tell.
  • The open-text base is 186 answers, not 548. Respondents who wrote nothing, or wrote only a short answer, are not represented, and themes are my coding of free text. Tagging is not exhaustive (for example, the performance and coding themes overlap with others).
  • The 20+ person segment has only 12 respondents. Any finding about the largest firms rests on very little data, and the head of sales' question about larger firms cannot be answered reliably from this survey.
  • Open-text answers are unprompted, whereas Q2 is prompted. The two measures should not be compared directly. For example, bulk MTD is mentioned by 22 paid respondents in open text but ranked first by 26.4% of paid respondents.

Next steps

  • decisionReframe the board paper: report the paid-practice ranking (feeds first) alongside the all-accounts view (AI first), and state that the all-accounts lead is not statistically significant against feeds. Do not describe the survey as backing AI categorisation.
  • researchDiagnose feed failures on paid accounts: classify each broken or late connection by cause (bank error, consent expiry, client not re-authorising, Ledgerline sync error) and measure how long each takes to be noticed.
  • experimentBuild the feed-reliability work with proactive alerts when consent is near expiry or a feed stops, and clearer re-authorisation prompts for clients, alongside the sync fixes the diagnosis supports.
  • experimentRun a priced pre-order or fake-door test for the AI categorisation add-on at GBP 8 per user per month, shown to free and paid accounts, and confirm whether free accounts can buy it at all.
  • researchRe-survey paid practices with quotas by practice size (2-9, 10-19, 20+) and record seat counts, with weighting to the 1,230 paid accounts. Include a question on document collection and on the reasons practices would leave.
  • researchCost bulk MTD submission as the runner-up, and check with the head of sales through win/loss interviews with 10 larger practices whether document collection or VAT batching drives decisions.
  • monitorTrack monthly the share of open-text and support mentions of feed failures, manual coding and one-at-a-time VAT submission, and the paid retention rate.

Open-text themes it coded

Bank feeds break, disconnect, expire or arrive late 72Manual categorisation of bank transactions 56Bulk MTD submission not possible: VAT returns filed one client at a time 25Chasing clients for documents, statements or receipts 15Slow performance (loading, search, reports) 26No complaint or nothing slows the practice down 25

The survey it planned

5 screening questions and 8 questions, as the model wrote them.

  1. S1

    Which of these best describes your role at your practice?

    One answer
    • I make the decisions about our software
    • I influence the decisions about our software
    • I do not make or influence decisions about our software

    Continues if I make the decisions about our software; I influence the decisions about our software

  2. S2

    Which of these best describes your practice?

    One answer
    • UK accountancy practice
    • UK bookkeeping practice
    • Both accountancy and bookkeeping
    • Other

    Continues if UK accountancy practice; UK bookkeeping practice; Both accountancy and bookkeeping

  3. S3

    Which Ledgerline plan does your practice use?

    One answer
    • Free plan
    • Paid plan
    • We do not use Ledgerline

    Continues if Free plan; Paid plan

  4. S4

    How many people work at your practice, including you?

    One answer
    • Just me (sole practitioner)
    • 2 to 9 people
    • 10 or more people

    Continues if Just me (sole practitioner); 2 to 9 people; 10 or more people

  5. S5

    Have you already started or completed this survey?

    One answer
    • No
    • Yes

    Continues if No

  6. Q1

    Which of these slow you down when you use Ledgerline today? Select all that apply.

    Any that apply
    • Bank feeds breaking or needing to be reconnected
    • Checking and correcting bank transaction categories by hand
    • Chasing clients for documents and information
    • Preparing and submitting Making Tax Digital returns one client at a time
    • Keeping track of client deadlines
    • Finding or updating client records
    • Nothing slows us down
    • Something else (please describe)
  7. Q2

    In a typical week, roughly how many hours does your practice spend working in Ledgerline? Please enter a whole number of hours.

    Number
  8. Q3

    In the last three months, how often have bank feed connections in Ledgerline stopped working and needed fixing?

    One answer
    • Never
    • Once or twice
    • Several times a month
    • Weekly or more often
    • We do not use bank feeds in Ledgerline
    Judges: ambiguous

    The options mix counts over three months with monthly and weekly frequencies, leaving some experiences without a clear answer—for example, three failures across the entire three-month period.

    The options mix two reference frames. 'Once or twice' is a count across the three months, while 'Several times a month' and 'Weekly' are rates. This leaves a gap: a practice with about 3 to 5 breaks in three months has no option that fits.

  9. Q4

    Rank these possible Ledgerline improvements from most useful to least useful for your practice. Put the most useful one first.

    Ranking
    • More reliable bank feeds, with fewer broken connections
    • AI categorisation of bank transactions, with suggested categories to review
    • A client portal, where clients upload the documents you request
    • Bulk Making Tax Digital submissions, for several clients at once
  10. Q5

    If Ledgerline could make only one of these improvements in the next six months, which would be most useful to your practice?

    One answer
    • More reliable bank feeds, with fewer broken connections
    • AI categorisation of bank transactions, with suggested categories to review
    • A client portal, where clients upload the documents you request
    • Bulk Making Tax Digital submissions, for several clients at once
    • None of these would help much
  11. Q6

    Suppose Ledgerline offered AI categorisation of bank transactions as an optional add-on. It would suggest a category for each transaction, and you would review and confirm them before they are saved. What is the most your practice would pay for this add-on each month?

    One answer
    • We would not pay for it
    • Up to £5 per user per month
    • £5 to £10 per user per month
    • £10 to £20 per user per month
    • More than £20 per user per month
    Judges: ambiguous

    The stem asks for the practice's maximum monthly payment, while the priced options are per user per month. The price bands also overlap at £5 and £10, so respondents with those maximum amounts have more than one applicable answer.

    The price options overlap (£5, £10, and £20 appear in multiple options), making it ambiguous which option a respondent should select if their maximum amount falls exactly on one of those boundaries.

  12. Q7

    To show you are reading carefully, please select 'Somewhat agree' for this question.

    Scale

    Scale 1-5: Strongly disagree to Strongly agree

  13. Q8

    What else slows your practice down in Ledgerline? Please describe it in a sentence or two.

    Open text

    Routing Shown only if Q1 is: Something else (please describe)

Sample plan

Population: the 8,200 Ledgerline practice accounts (about 6,970 on the free plan and 1,230 on paid plans). Invite the admin of every account by email, and show an in-app banner at login to all users. Each invitation carries a unique practice code so that only one response per practice is accepted; the screener S5 blocks repeat starts. Screen out anyone who does not make or influence software decisions, is not a UK accountancy or bookkeeping practice, or does not use Ledgerline. Quotas: each of the three size bands (sole practitioner, 2 to 9 staff, 10 or more staff) has a target of at least 384 completes, which gives about plus or minus 5 percentage points at 95% confidence for a 50% result, before any finite population correction. Within each size band, aim for a minimum of about 100 paid-plan completes so that the paid results are readable, and let free-plan completes fill the rest. Closing the field early in a band is allowed once its target is met. Total completes: about 1,150 to 1,200 if every band hits its target, not the 550 in the brief. At 550 completes, split evenly, each band would give about plus or minus 7 percentage points, so the precision requirement cannot be met at that size. Check the actual band sizes in the Ledgerline database first: where a band has few practices, the finite population correction reduces the number needed, and the 10 or more band may be small enough to census. Randomise the order of the four improvement options in Q4 and Q5 for each respondent, and keep the descriptions equally detailed. Weight the overall results to the practice population by size band and plan so that headline figures are representative. Report size-band results unweighted within band.