Benchmarks / SurveyBench / What Ledgerline's accountancy practices want built next / Mistral Large 4

Measured by Spring Prompt

Mistral Large 4: What Ledgerline's accountancy practices want built next

The decision: which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

Verdict
✗ Not sound
Research score
83 of 100
Analysis rating
815
Head to head, this task
won 6 of 15

Why it is not sound

The analysis

  • Reported a material finding the data does not support
  • Missed a trap: 28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • Proposed a mistaken next step: Forecast AI add-on revenue from the share who said they would pay £8 per user per month

Findings the judges found unsupported

  • AI categorisation is deeply polarising and would likely underperform commercially as an £8 add-on because paid users actively reject it, despite free users loving it.Ranking AI last establishes low relative priority, not active rejection or likely commercial underperformance. Hypothetical willingness to pay does not establish actual demand, and there is no investment-cost evidence to establish commercial viability.
  • Willingness to pay for AI categorisation at £8/user/month is weak among the paying customer base, suggesting limited add-on revenue.The estimate of 208 buying practices and £20k–£100k ARR converts hypothetical stated willingness to pay directly into purchasing behaviour. Neither actual conversion nor the assumed number of purchasing users is established.
  • The survey sample over-represents paid and highly engaged users, so 'total' figures skew toward the revenue base—appropriate for this decision, but free-user metrics should not be extrapolated to the entire free population.The sample demonstrably over-represents paid and frequent users, but there is no AI-preference breakdown by engagement showing that engaged free users inflate AI's first-place share. Reweighting by the known account plan mix would increase, not reduce, AI's first-place lead.
  • Manual bank categorisation is the dominant pain point for sole practitioners and free users, suggesting a potential lead-generation or freemium conversion strategy rather than a paid add-on.The supplied answers contain 56 mentions of manual categorisation, counting answers that also mention other themes, rather than the reported 54.
  • Bank feed reliability is the single most frequently cited pain point in qualitative responses, corroborating the quantitative ranking data.The supplied answers contain 74 bank-feed mentions and 56 manual-categorisation mentions, rather than 73 and 54. There are 25 MTD mentions, not the 26 stated in this finding. Bank feeds nevertheless remain the most frequently mentioned theme.

Scorecard

Numbers 11 of 11 right

QuestionIts answerTrue
What percentage of all respondents ranked AI categorisation of bank transactions first?✓37.8%37.8%
Among respondents on a paid plan, what percentage ranked more reliable bank feeds first?✓48.3%48.3%
What is the mean rank of more reliable bank feeds across all respondents (1 = highest priority, 4 = lowest), to two decimal places?✓1.991.99
What percentage of all respondents ranked AI categorisation of bank transactions last (4th)?✓42.2%42.2%
What percentage of all respondents said yes, their practice would pay £8 per user per month for AI categorisation?✓27.7%27.7%
What is the 95% margin of error, in percentage points, on the unweighted share of all respondents who ranked AI categorisation first?✓4.1 points4.06 points
Is the difference in the share ranking AI categorisation first between practices of 2-9 people and practices of 10-19 people statistically significant at the 95% level? (1 for yes, 0 for no)✓NoNo
How many of the open-text answers mention bank feeds breaking, disconnecting or arriving late?✓7374
How many of the open-text answers from respondents on a paid plan mention chasing clients for documents?✓1010
Ledgerline has 1,230 paid practices. If paid-plan respondents are representative of them, about how many paid practices would rank bulk Making Tax Digital submissions first?✓325324
What percentage of the 8,200 practice accounts invited completed the survey?✓6.7%6.7%

Traps in the data 6 of 7 handled

  • ✓AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • ✓Free-plan users are 46% of respondents but bring in no revenue; paid practices (15% of accounts, all of the revenue) put bank feeds first (48%) and bulk MTD (26%) ahead of AI (10%). The AI first-choice lead comes from free users (70% of whom rank it first). Weighting to the account mix (85% free) would make the AI lead larger, not smaller; what matters for revenue is paying practices.
  • ✗28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • ✓Every one of the 12 practices with 20 or more people ranked the client portal first.
  • ✓69% of completes came from an in-app banner seen only by users who logged in; 52% of banner respondents use Ledgerline on 5 or more days a week against 19% of email respondents and 18% of all accounts. The sample over-represents heavy users and the response rate is 6.7%.
  • ✓Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.
  • ✓The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

The analysis

Recommendation

Build more reliable bank feeds (fewer broken connections) next half-year. It is the top priority for the paying customer base (48.3% rank first, mean rank 1.76), has the lowest 'ranked last' rate (7.5%), and is the most cited pain point in open-text (73 mentions). Do not build AI categorisation as the next major initiative: it is polarising, rejected by 65.5% of paid users, and only 16.9% of paid users would pay £8/month for it—insufficient to justify the investment. Instead, treat AI categorisation as a potential freemium feature for free/sole practitioners (where 70.2% rank it first) to drive conversion, not as a paid add-on. Fast-follow with bulk MTD submissions, which is the clear second choice for paid practices (26.4% first, mean rank 2.23) and particularly valued by 10-19 person firms. Deprioritise the client portal until stronger evidence emerges from a larger sample of enterprise practices.

What would change it

The recommendation would change if: (1) a pricing experiment with real payment showed paid-user willingness to pay for AI categorisation exceeds 40% (vs 16.9% stated); (2) a larger sample of 50+ practices with 20+ employees showed statistically robust preference for the client portal (current n=12); (3) engineering assessment found bank feed reliability issues are technically intractable or prohibitively expensive to fix; or (4) churn analysis showed that broken bank feeds are not actually driving paid customers to leave, while manual categorisation is.

Findings

  1. More reliable bank feeds should be the next build, not AI categorisation. Paid customers—who generate 100% of revenue—rank bank feeds first by a wide margin, and it has the best average rank across all metrics.

    High confidenceBase 296

    48.3% of paid-plan respondents ranked bank feeds first vs only 10.1% for AI. Mean rank among paid users: bank feeds 1.76 (best) vs AI 3.35 (worst). Overall mean rank: bank feeds 1.99 vs AI 2.57. Bank feeds also has the lowest 'ranked last' rate at 7.5% (vs 42.2% for AI).

  2. AI categorisation is deeply polarising and would likely underperform commercially as an £8 add-on because paid users actively reject it, despite free users loving it.

    High confidenceBase 548

    42.2% of all respondents rank AI last—the highest 'last' rate of any option. Among paid plans, 65.5% rank it last and only 16.9% would pay £8/user/month. The 37.8% first-place figure is driven almost entirely by free-plan users (70.2% rank first), who are not the revenue base.

  3. The client portal should not be prioritised to win larger firms; the evidence that 'biggest practices want it' rests on a statistically meaningless sample of 12 respondents.

    High confidenceBase 118

    Only 12 respondents work at practices of 20+ people (100% ranked portal first, but n=12 is unreliable). Among 10-19 person practices (n=106), the portal is the least popular option: only 6.6% rank it first and 27.4% rank it last. For larger paid firms, bulk MTD (30.2% first among 10-19) and bank feeds are the clear priorities.

  4. Bulk Making Tax Digital submissions are the clear second priority for paid practices and should be fast-followed after bank feeds.

    High confidenceBase 296

    26.4% of paid respondents rank bulk MTD first (second only to bank feeds). Mean rank among paid: 2.23 (second best). 30.2% of 10-19 person practices rank it first. Open-text answers explicitly cite single-client MTD filing as a bottleneck at quarter-end.

  5. Willingness to pay for AI categorisation at £8/user/month is weak among the paying customer base, suggesting limited add-on revenue.

    Medium confidenceBase 296

    Only 16.9% of paid-plan respondents said yes (vs 40.5% of free users). 'Not sure' is 25.7% among paid users, indicating low conviction. Applied to 1,230 paid practices, this suggests roughly 208 buying practices; without user-count data, estimated ARR is approximately £20k-£100k (conservative) assuming 1-5 users per buying practice.

  6. The survey sample over-represents paid and highly engaged users, so 'total' figures skew toward the revenue base—appropriate for this decision, but free-user metrics should not be extrapolated to the entire free population.

    High confidenceBase 548

    54% of respondents are on paid plans vs 15% of the 8,200-account population. 41% use Ledgerline 5+ days/week vs 18% of the population. Response rate is 6.7%. This bias means the 'total' AI first-place vote (37.8%) is inflated by engaged free users.

  7. Manual bank categorisation is the dominant pain point for sole practitioners and free users, suggesting a potential lead-generation or freemium conversion strategy rather than a paid add-on.

    Medium confidenceBase 252

    70.2% of free-plan respondents rank AI categorisation first; 60.9% of 'Just me' practices rank it first. 40.5% of free users say they would pay £8/month. Open-text shows 54 mentions of manual coding as a time sink.

  8. Bank feed reliability is the single most frequently cited pain point in qualitative responses, corroborating the quantitative ranking data.

    High confidenceBase 186

    73 of 186 open-text answers mention feeds breaking, disconnecting, or arriving late—more than any other issue (manual categorisation: 54; MTD: 26; document chasing: 15; performance: 26).

Caveats

  • Self-selection bias: respondents were invited via in-app banner and email; 41% use Ledgerline 5+ days/week vs 18% of the population, and 54% are on paid plans vs 15% of the population. Results reflect engaged, paying users more than the total base.
  • Stated vs revealed preference: Q3 asks whether practices 'would pay' £8/month, not actual purchase behaviour. Willingness-to-pay in surveys typically overstates real conversion.
  • Tiny subgroups: practices of 20+ people have n=12; the 100% first-place vote for the client portal in this group is statistically meaningless and should not drive strategy.
  • Ranking data: first-place votes can be misleading when preferences are polarised. AI categorisation's 37.8% first-place share masks the fact that 42.2% rank it last; mean rank (1.99 for bank feeds vs 2.57 for AI) is a more robust measure of overall priority.
  • Unweighted data: tables are unweighted, so 'total' columns reflect the sample composition (54% paid) rather than the population (15% paid). Paid-subgroup figures are the relevant ones for revenue decisions.
  • Open-text coverage: only 186 of 548 completers (34%) provided open-text answers; theme frequencies may not represent the full sample.
  • Timing: fielded September 2026; sentiment may shift, particularly around January self-assessment and VAT deadlines.

Next steps

  • decisionApprove and build more reliable bank feeds: reduce disconnection frequency, add proactive monitoring/alerts when feeds go silent, and improve re-authorisation UX.
  • experimentRun a pricing experiment (van Westendorp or conjoint) with 200+ paid-plan practices to measure real willingness-to-pay for AI categorisation at £8 vs £5 vs £12.
  • researchConduct 20-30 in-depth interviews with paid practices of 10+ people to validate bulk MTD demand and explore client portal needs.
  • experimentPrototype bulk MTD submissions and test with 5-10 large paid practices (10+ people).
  • researchInvestigate free-user AI categorisation as a conversion lever: survey free users on whether AI categorisation would make them upgrade to paid, or offer as a 'freemium' feature to increase engagement.
  • monitorMonitor bank feed health metrics post-launch and correlate with churn/usage.

Open-text themes it coded

Bank feeds breaking, disconnecting, expiring or arriving late 73Manual categorisation/coding of bank transactions 54MTD submissions one-by-one / no bulk filing 25Chasing clients for documents, records or paperwork 15Slow performance, loading times or app lag 26No issues / positive sentiment 24

The survey it planned

3 screening questions and 8 questions, as the model wrote them.

  1. S1

    Is your practice based in the United Kingdom?

    One answer
    • Yes
    • No

    Continues if Yes

  2. S2

    Does your practice currently use Ledgerline?

    One answer
    • Yes
    • No

    Continues if Yes

  3. S3

    Do you make or influence decisions about which software your practice uses?

    One answer
    • Yes
    • No

    Continues if Yes

  4. Q1

    How many people work in your practice, including yourself?

    One answer
    • Just me (sole practitioner)
    • 2 to 9 people
    • 10 or more people
  5. Q2

    Which Ledgerline plan does your practice currently use?

    One answer
    • Free plan (one user)
    • Paid plan (charged per user per month)
    • Not sure
  6. Q3

    In a typical week, how often do you or your team actively use Ledgerline?

    One answer
    • Every day
    • Several times a week
    • About once a week
    • A few times a month
    • Rarely
  7. Q4

    What tasks or issues most slow you or your team down when using Ledgerline today? Please select up to three.

    Any that apply
    • Bank feeds disconnecting or failing to sync
    • Manually categorising bank transactions
    • Chasing clients for documents or information
    • Submitting Making Tax Digital (MTD) or VAT returns
    • Managing client deadlines and tasks
    • Something else (please specify)
    • Nothing significantly slows us down
  8. Q5

    Ledgerline is planning one major improvement in the next 6 months. Please rank these four options from 1 (most important for your practice) to 4 (least important).

    Ranking
    • AI categorisation of bank transactions (automatically suggesting categories for transactions)
    • Bulk Making Tax Digital submissions (submit MTD returns for multiple clients at once)
    • Client portal for requesting documents (clients upload documents directly)
    • More reliable bank feeds (fewer broken connections)
  9. Q6

    If Ledgerline offered AI categorisation of bank transactions as a paid add-on, how likely would your practice be to purchase it?

    Scale

    Scale 1-5: Very unlikely to Very likely

  10. Q7

    In the last 3 months, how often have you experienced broken bank feed connections in Ledgerline?

    One answer
    • Several times a week or more
    • About once a week
    • A few times
    • Once or twice
    • Never
    Judges: ambiguous

    The frequency options are inconsistent. 'A few times' and 'Once or twice' are undefined over a 3-month period, and there is a gap between 'About once a week' and 'A few times', so the categories overlap or leave holes.

    Assumes the practice uses bank feeds; 'Never' is ambiguous as it could mean they use bank feeds and they never break, or it could mean they do not use bank feeds at all.

  11. Q8

    Which of the following best describes your role in the practice?

    One answer
    • Owner / Partner / Director
    • Practice Manager
    • Bookkeeper / Accountant
    • Administrator / Other staff
Sample plan

Invite the admin of every UK Ledgerline practice account (n=8,200) by email, and display an in-app banner to all logged-in users. Target 600 completed responses (within the 'about 550' constraint, allowing for quality checks). To achieve ±5 percentage points precision at 95% confidence for each practice size group (sole practitioners, 2-9 staff, 10+ staff), 385 completes per group are required (total 1,155). Given the 600 target, we will apply soft quotas of 200 per practice size group, yielding ±6.9pp precision at 95% confidence per group; if strict ±5pp is required for all groups, the total must increase to 1,155. Use unique per-practice survey links (admin email + in-app token tied to practice ID) and deduplicate on practice ID to ensure one response per practice.