Benchmarks / SurveyBench / What Ledgerline's accountancy practices want built next / Gemini 3.5 Flash Lite

Measured by Spring Prompt

Gemini 3.5 Flash Lite: What Ledgerline's accountancy practices want built next

The decision: which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

Verdict
✗ Not sound
Research score
75 of 100
Analysis rating
662
Head to head, this task
won 4 of 10

Why it is not sound

The survey

  • Ignored the precision the brief asked for, which the sample cannot deliver
  • A double-barrelled question: Q3

The analysis

  • Got at least one set number wrong
  • Missed a trap: AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • Missed a trap: 28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.

Scorecard

Numbers 9 of 11 right

QuestionIts answerTrue
What percentage of all respondents ranked AI categorisation of bank transactions first?✓37.8%37.8%
Among respondents on a paid plan, what percentage ranked more reliable bank feeds first?✓48.3%48.3%
What is the mean rank of more reliable bank feeds across all respondents (1 = highest priority, 4 = lowest), to two decimal places?✓1.991.99
What percentage of all respondents ranked AI categorisation of bank transactions last (4th)?✓42.2%42.2%
What percentage of all respondents said yes, their practice would pay £8 per user per month for AI categorisation?✓27.7%27.7%
What is the 95% margin of error, in percentage points, on the unweighted share of all respondents who ranked AI categorisation first?✓4.1 points4.06 points
Is the difference in the share ranking AI categorisation first between practices of 2-9 people and practices of 10-19 people statistically significant at the 95% level? (1 for yes, 0 for no)✓NoNo
How many of the open-text answers mention bank feeds breaking, disconnecting or arriving late?✗5874
How many of the open-text answers from respondents on a paid plan mention chasing clients for documents?✗610
Ledgerline has 1,230 paid practices. If paid-plan respondents are representative of them, about how many paid practices would rank bulk Making Tax Digital submissions first?✓325324
What percentage of the 8,200 practice accounts invited completed the survey?✓6.7%6.7%

Traps in the data 5 of 7 handled

  • ✗AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • ✓Free-plan users are 46% of respondents but bring in no revenue; paid practices (15% of accounts, all of the revenue) put bank feeds first (48%) and bulk MTD (26%) ahead of AI (10%). The AI first-choice lead comes from free users (70% of whom rank it first). Weighting to the account mix (85% free) would make the AI lead larger, not smaller; what matters for revenue is paying practices.
  • ✗28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • ✓Every one of the 12 practices with 20 or more people ranked the client portal first.
  • ✓69% of completes came from an in-app banner seen only by users who logged in; 52% of banner respondents use Ledgerline on 5 or more days a week against 19% of email respondents and 18% of all accounts. The sample over-represents heavy users and the response rate is 6.7%.
  • ✓Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.
  • ✓The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

The analysis

Recommendation

Do not build AI categorisation next. Build more reliable bank feeds (fewer broken connections), as this is the primary operational bottleneck and top priority for the paying subscriber base that generates Ledgerline's revenue.

What would change it

Definitive economic proof that building AI categorisation would drive massive free-to-paid conversion upgrades outweighing the churn risk and low satisfaction among existing paying users, or empirical telemetry showing bank feed stability has reached 99.9% reliability independently of software updates.

Findings

  1. AI categorisation of bank transactions is a polarising feature driven almost entirely by free-plan users and solo practices, rather than the paying customer base.

    High confidenceBase 548

    70.2% of free-plan users ranked AI categorisation first, compared to only 10.1% of paid-plan users. Conversely, 65.5% of paid-plan users ranked it last (4th).

  2. More reliable bank feeds are the top priority for paying practices and represent the core operational pain point across the user base.

    High confidenceBase 548

    48.3% of paid-plan users ranked bank feeds first, giving it a mean rank of 1.76. 58 open-text comments explicitly mentioned broken, disconnecting, or late bank feeds.

  3. The client portal is not the primary desire for larger practices, contrary to the sales team's assumption.

    Medium confidenceBase 548

    Among practices with 10-19 people, only 6.6% ranked the client portal first (against 49.1% for bank feeds and 30.2% for bulk MTD). Even among 20+ person practices (base 12), the sample is too small to substantiate the sales hypothesis as a general rule.

  4. Demand for bulk Making Tax Digital (MTD) submissions scales directly with practice size among paying users.

    High confidenceBase 296

    26.4% of paid users overall ranked bulk MTD first, rising to 30.2% for practices of 10-19 people, making it a critical feature for medium-to-large paying firms.

Caveats

  • Survey respondents skew heavily toward daily users (52% of banner respondents use the software 5+ days a week vs 18% in the total account population), meaning power-user pain points are over-represented.
  • The overall response rate is 6.7% (548 completes out of 8,200 accounts), which introduces potential self-selection bias from users with strong positive or negative opinions.
  • Subgroups for very large practices (20+ people) contain extremely small bases (n=12), limiting generalisability for enterprise-tier firms.

Next steps

  • decision Brief the CTO that raw ranking totals conflate free users with revenue-generating practices, and present the segmented plan data showing paid-user preference for bank feeds.
  • decisionModel potential add-on uptake vs conversion impact with the Finance Director, noting that only 16.9% of paid users say they would pay £8/mo for AI categorisation.
  • researchReview sales pipeline data with the Head of Sales to test whether client portal requests are concentrated in prospects rather than established larger firms.
  • monitorConduct technical discovery on bank connection stability (e.g., OAuth token persistence, Open Banking API error rates) to scope the bank feed reliability project.

Open-text themes it coded

Manual bank transaction coding/categorisation 49Bank feeds breaking, disconnecting or arriving late 58Individual/non-bulk MTD VAT submissions slowing down workflows 25Chasing clients for documents, records and approvals 13App performance, lag and slow page loads 24

The survey it planned

2 screening questions and 8 questions, as the model wrote them.

  1. SCR1

    Do you make or heavily influence decisions about the software and tools used in your accountancy or bookkeeping practice?

    One answer
    • Yes, I am the sole decision-maker or part of the decision-making team
    • No, I do not influence software decisions

    Continues if Yes, I am the sole decision-maker or part of the decision-making team

  2. SCR2

    Which of the following best describes your accountancy or bookkeeping practice?

    One answer
    • Sole practitioner (no other staff)
    • Small practice with 2 to 9 staff
    • Medium or large practice with 10 or more staff
    • Not an accountancy or bookkeeping practice

    Continues if Sole practitioner (no other staff); Small practice with 2 to 9 staff; Medium or large practice with 10 or more staff

  3. Q1

    Which Ledgerline plan is your practice currently on?

    One answer
    • Free plan
    • Paid plan
    • Not sure
  4. Q2

    How often do members of your practice log into or use Ledgerline?

    One answer
    • Daily
    • 2-3 times a week
    • About once a week
    • Less than once a week
  5. Q3

    What are the biggest bottlenecks or tasks that slow your practice down while using Ledgerline? (Select all that apply)

    Any that apply
    • Reconciling bank transactions and fixing broken bank feeds
    • Chasing clients for documents and records
    • Preparing and submitting Making Tax Digital (MTD) returns one by one
    • Manually categorising routine transactions
    • Managing deadlines and task statuses
    • Other (please specify)
    • None, everything runs smoothly
    Judges: double-barrelled

    One response option combines reconciling bank transactions with fixing broken bank feeds. Selecting it does not establish which task is a bottleneck, although that distinction matters to the build decision.

    The first option joins two separate problems: 'Reconciling bank transactions' and 'fixing broken bank feeds'. A respondent bothered by only one must tick both or neither. This also blurs the link between the bank-feed improvement and the other options.

  6. Q4

    The product team has room to build one major new capability in the next half-year. Please rank the following four options from 1 (what you want most) to 4 (what you want least):

    Ranking
    • AI categorisation of bank transactions
    • More reliable bank feeds with fewer broken connections
    • A client portal for requesting documents
    • Bulk Making Tax Digital (MTD) submissions
  7. Q5

    How much do broken bank-feed connections impact your practice's daily efficiency?

    Scale

    Scale 1-5: No impact at all to Severe impact / major frustration

  8. Q6

    If Ledgerline offered AI categorisation of bank transactions as an optional paid add-on feature, would your practice be willing to pay an extra monthly fee for it?

    One answer
    • Yes, definitely
    • Maybe, depending on the price
    • No, I would expect this to be included in our existing plan
    • No, we have no interest in AI categorisation
  9. Q7

    What is the maximum additional monthly amount per user your practice would be willing to pay for AI categorisation of bank transactions?

    One answer
    • Under £5 per user/month
    • £5 to £10 per user/month
    • £11 to £20 per user/month
    • More than £20 per user/month
    • Zero / Not willing to pay extra

    Routing Shown only if Q6 is: Yes, definitely; Maybe, depending on the price

  10. Q8

    Is there anything else you would like the Ledgerline product team to consider when planning updates for the next half-year?

    Open text
Sample plan

The survey will be distributed via an in-app banner upon login and an email invitation sent to the designated admin/decision-maker of all 8,200 Ledgerline practice accounts (comprising both the 6,970 free plan accounts and 1,230 paid plan accounts). To achieve the target of approximately 550 completes and ensure sub-group precision within +/- 5 percentage points at a 95% confidence level for each practice size category, we will monitor incoming completes and apply soft quotas across practice size tiers (sole practitioners, 2-9 staff, 10+ staff) and plan types (free vs paid).