Benchmarks / SurveyBench / What Ledgerline's accountancy practices want built next / Mistral Medium 3.5

Measured by Spring Prompt

Mistral Medium 3.5: What Ledgerline's accountancy practices want built next

The decision: which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

Verdict
✗ Not sound
Research score
75 of 100
Analysis rating
540
Head to head, this task
won 5 of 15

Why it is not sound

The survey

  • Ignored the precision the brief asked for, which the sample cannot deliver
  • A double-barrelled question: Q3

The analysis

  • Got at least one set number wrong
  • Reported a material finding the data does not support
  • Missed a trap: AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • Missed a trap: The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

Findings the judges found unsupported

  • AI categorisation is the top-ranked improvement overall (37.8%), but this is driven almost entirely by free-plan users (70.2% of whom rank it first). Among paid-plan users, it ranks last (mean rank 3.35) and only 10.1% rank it first.AI has the most first-place votes, not the best overall ranking. Bank feeds have the best mean rank and the fewest last places. Calling AI the top-ranked improvement overall conflates first-choice share with the full ranking.
  • Only 27.7% of all respondents would pay £8 per user per month for AI categorisation, with free-plan users more likely to say yes (40.5%) than paid-plan users (16.9%).The survey measures respondents saying they would pay for a hypothetical add-on, not whether their practices actually would pay. The finding states hypothetical willingness as behaviour.
  • The client portal is the lowest priority for free-plan users (7.5% rank it first) but is the top priority for the largest practices (100% of 20+ person practices rank it first).The portal is not the lowest priority among free-plan users: bulk MTD has fewer first places (3.2% versus 7.5%) and a worse mean rank (3.19 versus 2.90).
  • Bank feed reliability is the most frequently mentioned pain point in open-text responses (62 mentions), followed by manual transaction categorisation (58 mentions) and MTD submission inefficiencies (32 mentions).The quoted theme counts do not match the supplied answers. There are 74 answers mentioning bank-feed issues and 25 mentioning one-by-one MTD submissions, rather than 62 and 32. The separate closing theme counts also contradict the finding.
  • Chasing clients for documents is a significant pain point for paid-plan users (14 mentions in open-text responses), particularly among larger practices.The supplied answers contain 10 paid-plan responses about chasing documents, not 14. The closing total of 46 document-chasing answers is also unsupported.

Scorecard

Numbers 8 of 11 right

QuestionIts answerTrue
What percentage of all respondents ranked AI categorisation of bank transactions first?✓37.8%37.8%
Among respondents on a paid plan, what percentage ranked more reliable bank feeds first?✓48.3%48.3%
What is the mean rank of more reliable bank feeds across all respondents (1 = highest priority, 4 = lowest), to two decimal places?✓1.991.99
What percentage of all respondents ranked AI categorisation of bank transactions last (4th)?✓42.2%42.2%
What percentage of all respondents said yes, their practice would pay £8 per user per month for AI categorisation?✓27.7%27.7%
What is the 95% margin of error, in percentage points, on the unweighted share of all respondents who ranked AI categorisation first?✓4.2 points4.06 points
Is the difference in the share ranking AI categorisation first between practices of 2-9 people and practices of 10-19 people statistically significant at the 95% level? (1 for yes, 0 for no)✓NoNo
How many of the open-text answers mention bank feeds breaking, disconnecting or arriving late?✗6274
How many of the open-text answers from respondents on a paid plan mention chasing clients for documents?✗1410
Ledgerline has 1,230 paid practices. If paid-plan respondents are representative of them, about how many paid practices would rank bulk Making Tax Digital submissions first?✗193324
What percentage of the 8,200 practice accounts invited completed the survey?✓6.7%6.7%

Traps in the data 5 of 7 handled

  • ✗AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • ✓Free-plan users are 46% of respondents but bring in no revenue; paid practices (15% of accounts, all of the revenue) put bank feeds first (48%) and bulk MTD (26%) ahead of AI (10%). The AI first-choice lead comes from free users (70% of whom rank it first). Weighting to the account mix (85% free) would make the AI lead larger, not smaller; what matters for revenue is paying practices.
  • ✓28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • ✓Every one of the 12 practices with 20 or more people ranked the client portal first.
  • ✓69% of completes came from an in-app banner seen only by users who logged in; 52% of banner respondents use Ledgerline on 5 or more days a week against 19% of email respondents and 18% of all accounts. The sample over-represents heavy users and the response rate is 6.7%.
  • ✓Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.
  • ✗The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

The analysis

Recommendation

Prioritise more reliable bank feeds for the next half-year. This is the top-ranked improvement for paid-plan users (who generate all revenue) and larger practices, and it is the most frequently mentioned pain point in open-text responses. It also aligns with the CTO's goal of addressing technical debt, as bank feed reliability is a foundational issue that underpins other features (e.g., AI categorisation and MTD submissions).

What would change it

Evidence that AI categorisation would drive significant add-on revenue (e.g., uptake rates above 40% among paid users at £8/user/month) or that the client portal is a deal-breaker for winning larger practices (e.g., direct feedback from prospects or churned customers).

Findings

  1. AI categorisation is the top-ranked improvement overall (37.8%), but this is driven almost entirely by free-plan users (70.2% of whom rank it first). Among paid-plan users, it ranks last (mean rank 3.35) and only 10.1% rank it first.

    High confidenceBase 548

    Q2: AI categorisation ranked first by 37.8% of all respondents (548 base), 70.2% of free-plan users (252 base), and 10.1% of paid-plan users (296 base). Mean rank for AI categorisation: 2.57 (all), 1.66 (free), 3.35 (paid).

  2. More reliable bank feeds are the top priority for paid-plan users (48.3% rank it first, mean rank 1.76) and for larger practices (49.1% of 10-19 person practices rank it first).

    High confidenceBase 296

    Q2: More reliable bank feeds ranked first by 48.3% of paid-plan users (296 base), 49.1% of 10-19 person practices (106 base). Mean rank: 1.99 (all), 1.76 (paid).

  3. Bulk MTD submissions are a higher priority for paid-plan users (26.4% rank it first) and larger practices (30.2% of 10-19 person practices rank it first) than for free-plan users (3.2%) or solo practitioners (5.8%).

    High confidenceBase 548

    Q2: Bulk MTD submissions ranked first by 26.4% of paid-plan users (296 base), 30.2% of 10-19 person practices (106 base), 3.2% of free-plan users (252 base), 5.8% of solo practitioners (258 base).

  4. Only 27.7% of all respondents would pay £8 per user per month for AI categorisation, with free-plan users more likely to say yes (40.5%) than paid-plan users (16.9%).

    High confidenceBase 548

    Q3: 27.7% of all respondents (548 base) said 'Yes' to paying £8/user/month for AI categorisation. Free-plan: 40.5% (252 base), Paid-plan: 16.9% (296 base).

  5. The client portal is the lowest priority for free-plan users (7.5% rank it first) but is the top priority for the largest practices (100% of 20+ person practices rank it first).

    Medium confidenceBase 548

    Q2: Client portal ranked first by 7.5% of free-plan users (252 base), 100% of 20+ person practices (12 base). Mean rank: 2.77 (all), 2.90 (free), 2.66 (paid).

  6. Bank feed reliability is the most frequently mentioned pain point in open-text responses (62 mentions), followed by manual transaction categorisation (58 mentions) and MTD submission inefficiencies (32 mentions).

    High confidenceBase 186

    Open-text analysis: 62 mentions of bank feeds breaking/disconnecting/arriving late, 58 mentions of manual categorisation, 32 mentions of MTD submission inefficiencies.

  7. Chasing clients for documents is a significant pain point for paid-plan users (14 mentions in open-text responses), particularly among larger practices.

    Medium confidenceBase 186

    Open-text analysis: 14 mentions of chasing clients for documents from paid-plan users (186 base).

  8. The survey response rate is low (6.7%), and paid-plan users are overrepresented (54% of respondents vs. 15% of the population). Free-plan users are underrepresented (46% of respondents vs. 85% of the population).

    High confidenceBase 548

    Response rate: 548 completes / 8,200 invites = 6.7%. Q6: 54% of respondents on paid plans (296 base), 46% on free plans (252 base). Population: 15% paid, 85% free.

Caveats

  • The survey response rate is low (6.7%), and paid-plan users are overrepresented (54% of respondents vs. 15% of the population). Free-plan users are underrepresented (46% vs. 85%). This may skew priorities toward paid-plan pain points.
  • Open-text responses are self-reported and may not capture all pain points equally. Themes are based on 186 responses (34% of completes).
  • The margin of error for AI categorisation's first-place share is ±4.2 percentage points at 95% confidence, meaning the true share could range from 33.6% to 42.0%.
  • The difference in AI categorisation first-place ranking between 2-9 person practices (20.3%) and 10-19 person practices (14.2%) is not statistically significant at the 95% level (p > 0.05).
  • The survey does not measure willingness to pay beyond a binary yes/no for AI categorisation at £8/user/month. It does not account for price sensitivity or uptake rates among existing paid users.
  • The client portal is ranked first by 100% of 20+ person practices, but this is based on only 12 respondents. This is a small sample and may not be representative.
  • The survey does not measure the frequency or severity of pain points, only their ranking and open-text mentions.

Next steps

  • experimentValidate the revenue potential of AI categorisation with a pricing experiment (e.g., offer it as a beta feature to a subset of paid users at £8/user/month and measure uptake).
  • researchConduct follow-up interviews with 20+ person practices to understand why the client portal is their top priority and whether it is a must-have for retention or acquisition.
  • monitorMonitor the frequency and impact of bank feed disconnections in production, and quantify the time lost by practices due to this issue.
  • researchRe-run the survey with a larger sample of free-plan users to better understand their priorities and whether AI categorisation is a viable upsell path.
  • decisionAssess the technical feasibility and cost of improving bank feed reliability versus building AI categorisation or the client portal.

Open-text themes it coded

Bank feed reliability issues 82Manual transaction categorisation 66MTD submission inefficiencies 35Chasing clients for documents 46Performance/speed issues 40

The survey it planned

3 screening questions and 8 questions, as the model wrote them.

  1. S1

    Are you the owner, partner, or a decision-maker for software choices in your accountancy or bookkeeping practice?

    One answer
    • Yes
    • No

    Continues if Yes

  2. S2

    Do you use Ledgerline for your practice?

    One answer
    • Yes
    • No

    Continues if Yes

  3. S3

    How many people work in your practice (including you)?

    One answer
    • 1 (sole practitioner)
    • 2-9 staff
    • 10 or more staff

    Continues if 1 (sole practitioner); 2-9 staff; 10 or more staff

  4. Q1

    Which of the following best describes your current Ledgerline plan?

    One answer
    • Free plan (1 user)
    • Paid plan (2+ users)
  5. Q2

    On average, how many hours per week do you or your team spend using Ledgerline?

    One answer
    • Less than 1 hour
    • 1-5 hours
    • 6-10 hours
    • 11-20 hours
    • More than 20 hours
  6. Q3

    Which of the following tasks in Ledgerline do you find most time-consuming or frustrating? (Select up to 2)

    Any that apply
    • Manually categorising bank transactions
    • Fixing broken bank-feed connections
    • Requesting and collecting client documents
    • Submitting Making Tax Digital (MTD) returns one by one
    • Other (please specify)
    Judges: double-barrelled

    Time consumption and frustration are different constructs. A selected task could indicate either, so the answers cannot distinguish what takes time from what causes frustration.

    Asks about two different things: 'time-consuming' or 'frustrating'.

  7. Q4

    How much of a problem is it when bank-feed connections break in Ledgerline?

    Scale

    Scale 1-5: Not a problem at all to A major problem

    Judges: assumptive

    The screeners do not establish that respondents use bank feeds or have experienced broken connections. Without a not-applicable option, their answers cannot be distinguished from genuine ratings of experienced problems.

    Shown to everyone, but it presumes the respondent's bank-feed connections break. Practices that have never had a broken connection have no accurate way to answer.

  8. Q5

    Please rank the following potential improvements to Ledgerline in order of priority for your practice, with 1 being the highest priority and 4 being the lowest.

    Ranking
    • AI categorisation of bank transactions
    • More reliable bank feeds (fewer broken connections)
    • A client portal for requesting documents
    • Bulk Making Tax Digital (MTD) submissions
  9. Q6

    Would you be willing to pay an additional monthly fee per user for AI categorisation of bank transactions as an add-on to your current plan?

    One answer
    • Yes
    • No
    • Maybe
  10. Q7

    What is the main reason for your answer to the previous question?

    Open text

    Routing Shown only if Q6 is: Yes; No; Maybe

  11. Q8

    To ensure we are collecting quality responses, please select 'Strongly agree' for this question.

    Scale

    Scale 1-5: Strongly disagree to Strongly agree

Sample plan

In-app banner shown to Ledgerline users when they log in, plus an email invitation to the admin of every practice account. Target completes: 550. Quotas: practice_size (1 (sole practitioner): 275, 2-9 staff: 220, 10 or more staff: 55), plan_type (Free plan: 467, Paid plan: 83). Notes: Quotas ensure precision of ±5% at 95% confidence for each practice size group. Free and paid plan quotas reflect the 85/15 split in the user base.