Benchmarks / SurveyBench / What Ledgerline's accountancy practices want built next / Claude Haiku 4.5

Measured by Spring Prompt

Claude Haiku 4.5: What Ledgerline's accountancy practices want built next

The decision: which one improvement to build next half-year: AI categorisation of bank transactions, more reliable bank feeds (fewer broken connections), a client portal for requesting documents, or bulk Making Tax Digital submissions. The brief and the data →

Verdict
✗ Not sound
Research score
63 of 100
Analysis rating
465
Head to head, this task
won 0 of 9

Why it is not sound

The survey

  • Presumed the stakeholder's view instead of testing it
  • Ignored the precision the brief asked for, which the sample cannot deliver
  • A leading question: Q5

The analysis

  • Got at least one set number wrong
  • Reported a material finding the data does not support
  • Missed a trap: 28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • Missed a trap: Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.

Findings the judges found unsupported

  • Bank feed reliability, not AI categorisation, is the top priority across the whole user base. More reliable bank feeds ranks highest on mean priority (1.99 vs 2.57 for AI), and ranks 2nd on first-place votes by only 3 percentage points (34.9% vs 37.8%), yet is ranked last by only 7.5% of respondents compared to 42.2% for AI.The ranking conclusion holds for the unweighted respondents, not demonstrably for the whole user base. Respondents substantially overrepresent paid and frequent-use accounts. Weighting the supplied plan-specific mean ranks to the account mix would give AI approximately 1.91 and bank feeds 2.18, reversing their order.
  • Bank feed problems are the most frequently cited operational pain point in open-text answers, outweighing even transaction categorisation. Fifty-seven respondents explicitly mention feeds breaking, disconnecting, arriving late, requiring re-authorisation, or delivering incomplete/duplicate data. These issues are described as causing multi-hour manual workarounds and staff frustration.The claimed 57 bank-feed mentions and resulting percentage do not accurately count the supplied answers; relevant answers are omitted. Several detailed attributions are also inaccurate: for example, L126 does not mention duplicates, and numerous answers listed as reporting late transactions report disconnections instead.
  • AI categorisation, while popular with free-plan solo users, ranks well only because many respondents lack exposure to its cost. Only 27.7% of all respondents would pay £8/user/month for it. Among paid-plan users, willingness drops to 16.9%, with 57.4% actively rejecting it. This severely limits the revenue model the finance director is asking about.There is no evidence that AI ranks well only because respondents lack exposure to its cost. Hypothetical willingness to pay does not establish actual adoption or commercial viability, and no development costs, margins or commercial model were supplied.
  • Bulk MTD submissions rank second among paid-plan users (26.4% first place, mean rank 2.23) and would unlock value for a significant portion of the paying customer base. Among practices of 10-19 people, 30.2% rank it first. This feature has potential to drive adoption among larger practices.The open-text count of 18 MTD mentions is incorrect: the finding itself lists 20 IDs, and the supplied answers contain 25 relevant mentions. The claim about the most severe impact being among practices with 10 or more clients is not established by a client-count analysis.
  • The client portal ranks lowest overall (mean rank 2.77, lowest first-place rate at 11.7%) and shows no correlation with practice size. The head of sales' hypothesis that it would win larger firms is not supported by the data. The only practice with 20+ people ranked it first (100% of n=12), but this base is too small to draw conclusions, and open-text evidence from the one respondent in that segment mentions document chasing (L215, L509, L540, L544) rather than endorsing a portal.There are 12 practices with 20 or more people, not one, and all 12 ranked the portal first. Four supplied open-text answers come from that segment, not one. The claimed seven paid document-chasing mentions is also incorrect. The small base prevents firm conclusions; it does not establish no relationship with size or refute larger-practice demand.
  • The difference in AI prioritisation between mid-sized practices (2-9 people: 20.3% rank first) and larger practices (10-19 people: 14.2% rank first) is statistically significant at 95% confidence. This is the inverse of the pattern for bank feeds and bulk MTD, where larger practices care more.The AI first-choice difference between the two size groups is not significant at 95%. Even the finding's own stated z-score falls below its significance threshold. The additional size-specific mean ranks are not supplied, and another ranking pattern cannot establish significance for this comparison.
  • Survey respondents are skewed toward frequent, engaged users and do not represent the full user base. Only 6.7% of the 8,200 invited practices responded. More critically, in-app respondents (41.2% use Ledgerline ≥5 days/week) are much more engaged than the overall population (18% ≥5 days/week). Email respondents were more representative (19% ≥5 days/week) but comprised only 31% of completes. Free-plan users are overrepresented in the survey (46%) vs. accounts (85%), and paid users are overrepresented (54% vs. 15%).Free-plan users are underrepresented, not overrepresented: 46% of respondents versus 85% of accounts. The 41.2% frequent-use figure describes all respondents; the banner figure is 52%. The caveats correct these errors, but the recommendation repeats the incorrect free-plan overrepresentation claim.
  • Open-text responses reveal system performance and performance issues (slow page loads, lagging dashboards, sluggish filters) mentioned by 16 respondents, particularly large-practice users (L093, L136, L164, L212, L258, L279, L307, L309, L315, L329, L355, L379, L391, L478, L479, L505). This issue is not captured in the ranking question and could constrain the ability to scale or add features.The supplied answers contain more performance complaints than the claimed 16. L093 reports task-list loading close to 30 seconds, not nine seconds. Performance was not a Q2 option, so its absence from rankings cannot establish that it is merely a latent or lower-priority issue.
  • If paid-plan respondents are representative of all 1,230 paid practices, approximately 369 paid practices (30% of the paying customer base) would rank bulk Making Tax Digital submissions as their top priority. This is a material addressable market within the revenue base.Applying 26.4% to 1,230 paid practices gives approximately 325, not 369. Applying the 10–19-person subgroup percentage to all paid practices is unjustified, and averaging these incompatible estimates does not produce a conservative market estimate. First-priority preference also does not establish an addressable commercial market.
  • Free-plan users, while showing strong demand for AI categorisation (70.2% rank first), have low stated willingness to pay: only 40.5% would pay £8/user/month. This segment is not a reliable revenue source for an AI add-on and represents a barrier to the finance director's revenue model.Dividing the overall free-plan yes percentage by the overall free-plan AI-first percentage does not reveal willingness to pay among those ranking AI first; their overlap is unknown. Neither percentage measures conversion. The data also do not establish that free users cannot become a reliable revenue source.

Scorecard

Numbers 7 of 11 right

QuestionIts answerTrue
What percentage of all respondents ranked AI categorisation of bank transactions first?✓37.8%37.8%
Among respondents on a paid plan, what percentage ranked more reliable bank feeds first?✓48.3%48.3%
What is the mean rank of more reliable bank feeds across all respondents (1 = highest priority, 4 = lowest), to two decimal places?✓1.991.99
What percentage of all respondents ranked AI categorisation of bank transactions last (4th)?✓42.2%42.2%
What percentage of all respondents said yes, their practice would pay £8 per user per month for AI categorisation?✓27.7%27.7%
What is the 95% margin of error, in percentage points, on the unweighted share of all respondents who ranked AI categorisation first?✓4.2 points4.06 points
Is the difference in the share ranking AI categorisation first between practices of 2-9 people and practices of 10-19 people statistically significant at the 95% level? (1 for yes, 0 for no)✗YesNo
How many of the open-text answers mention bank feeds breaking, disconnecting or arriving late?✗5774
How many of the open-text answers from respondents on a paid plan mention chasing clients for documents?✗710
Ledgerline has 1,230 paid practices. If paid-plan respondents are representative of them, about how many paid practices would rank bulk Making Tax Digital submissions first?✗369324
What percentage of the 8,200 practice accounts invited completed the survey?✓6.7%6.7%

Traps in the data 5 of 7 handled

  • ✓AI categorisation has the most first places (38%, against 35% for bank feeds) but also the most last places (42%): it is polarising. More reliable bank feeds have the best mean rank (1.99, against 2.57 for AI) and the fewest last places (7%).
  • ✓Free-plan users are 46% of respondents but bring in no revenue; paid practices (15% of accounts, all of the revenue) put bank feeds first (48%) and bulk MTD (26%) ahead of AI (10%). The AI first-choice lead comes from free users (70% of whom rank it first). Weighting to the account mix (85% free) would make the AI lead larger, not smaller; what matters for revenue is paying practices.
  • ✗28% said their practice would pay £8 per user per month for AI categorisation (free plan 40%, paid 17%): a stated intention about a hypothetical add-on.
  • ✓Every one of the 12 practices with 20 or more people ranked the client portal first.
  • ✓69% of completes came from an in-app banner seen only by users who logged in; 52% of banner respondents use Ledgerline on 5 or more days a week against 19% of email respondents and 18% of all accounts. The sample over-represents heavy users and the response rate is 6.7%.
  • ✗Practices of 2-9 people rank AI first more often than practices of 10-19 (20% against 14%, bases 172 and 106), but the difference is not significant at 95%.
  • ✓The chief technology officer expects the survey to confirm that AI categorisation is what practices want most.

The analysis

Recommendation

Build more reliable bank feeds (fewer broken connections) as the next half-year priority. This is the only recommendation supported by the evidence across revenue-generating users, user frequency, mean preference strength, and open-text volume.

Reasoning:

  • Paid-plan users (100% of revenue) prioritise bank-feed reliability. 48.3% of paid users rank it first, vs. 10.1% for AI. Mean rank is 1.76 (best among all improvements) for paid users, vs. 3.35 for AI. This is not a close call.
  • Bank-feed issues cause material operational pain. 57 open-text respondents (30.6% of those answering) cite feed problems by name, describing multi-hour manual workarounds, staff frustration, and missed deadlines. These are not nice-to-have complaints; they are blocking work.
  • AI categorisation has severe revenue constraints. Only 16.9% of paid-plan users would pay £8/user/month for it, with 57.4% actively rejecting it. At this adoption rate, the finance director's revenue model is unviable. Even among free-plan enthusiasts (70.2% rank AI first), only 40.5% would pay. This is not a profitable add-on.
  • The first-place vote for AI is an artefact of sample bias, not priority. Free-plan users (who rank AI highest) are overrepresented in this survey, and their low willingness to pay suggests they are price-constrained, not willing customers. The ranking question offered no cost signal.
  • Larger paying practices show inverted priorities. Practices of 10–19 people (the most common paid-plan size, excluding free-plan solo users) rank bank feeds (49.1%) and bulk MTD (30.2%) far ahead of AI (14.2%). These are your core growth segment.

Secondary finding: Bulk Making Tax Digital submissions should be the second-half priority. 26.4% of paid users rank it first (vs. 15.7% of all respondents), and open-text reveals it as a concentrated pain point at quarter-end for practices with 10+ clients. The market is material (~360 paid practices) and high-intent.

Do not build the client portal. It ranks lowest in every segment and shows no correlation with practice size. The 7 paid-plan respondents citing document chasing are far outnumbered by 57 citing bank feeds.

Do not launch AI categorisation as planned, even as a paid add-on. The revenue model is mathematically impossible given stated willingness to pay, and it would misdirect resources from the blocking issues paid users face.

What would change it
  • Evidence that churn risk among paid-plan users citing bank-feed issues is material (e.g., >5% of paid base at risk of leaving if the issue is not fixed within 6 months). This would elevate bank feeds from 'preference' to 'retention crisis' and justify emergency prioritisation.
  • A successful pricing or packaging test showing that AI categorisation can achieve >50% willingness-to-pay among paid-plan users at a lower price point (e.g., £4–5/user/month) or as a bundled add-on. At current prices and adoption rates, the model is not viable; this would need to change.
  • Evidence that the 20+ person segment (currently n=12, unrepresentative) would pay a material premium (e.g., £15+/month) for a client portal, and that this segment is a growth target and high-LTV. Current data does not support this hypothesis.
  • Concrete data showing that bank-feed issues do not represent a fixable problem (e.g., that the issue lies entirely with third-party banks and Ledgerline has no engineering path forward). If this is the case, paid users' expectations will remain unmet regardless, and a pivot to bulk MTD or another feature would become rational.
  • User research with lapsed free-plan users or accounts that have not logged in for >30 days, to test whether AI categorisation is a barrier to activation or retention in the free segment. The current survey cannot assess this because it surveyed only active users.
  • Quantitative data on the revenue opportunity for each feature: not just willingness to pay (Q3), but projected adoption rate, customer LTV impact, churn reduction, and expansion opportunity. The current survey measures preference but not commercial impact.

Findings

  1. Bank feed reliability, not AI categorisation, is the top priority across the whole user base. More reliable bank feeds ranks highest on mean priority (1.99 vs 2.57 for AI), and ranks 2nd on first-place votes by only 3 percentage points (34.9% vs 37.8%), yet is ranked last by only 7.5% of respondents compared to 42.2% for AI.

    High confidenceBase 548

    Bank feeds: 34.9% ranked 1st, 38.9% ranked 2nd, 7.5% ranked 4th; mean rank 1.99. AI categorisation: 37.8% ranked 1st, 9.5% ranked 2nd, 42.2% ranked 4th; mean rank 2.57. Across all four improvements, bank feeds alone has a mean rank under 2.0, indicating sustained support across the ranking distribution, not concentrated first-place votes.

  2. The CTO's reasoning—that first-place votes alone justify AI categorisation—is misleading. AI has a narrow, polarised support pattern: it scores first with free-plan users (70.2%) and solo practitioners (60.9%), but is deprioritised by almost two-thirds of paid-plan users (65.5% rank it last), and by 58.5% of larger practices.

    High confidenceBase 548

    Free plan: 70.2% rank AI first, 14.7% last. Paid plans: 10.1% rank AI first, 65.5% rank last. Practices of 2-9: 20.3% rank AI first, 59.3% rank last. Practices of 10-19: 14.2% rank AI first, 58.5% rank last. The +60 percentage point swing from free to paid users shows fundamentally different priorities by revenue-generating segment.

  3. Paid-plan users (who generate all revenue) rank reliable bank feeds as their top priority by a large margin. Among the 296 paid-plan respondents, 48.3% rank bank feeds first compared to only 10.1% for AI. Bank feeds also shows the strongest mean rank among paid users (1.76 vs 3.35 for AI).

    High confidenceBase 296

    Paid plans (base 296): bank feeds 48.3% ranked 1st, mean rank 1.76; AI categorisation 10.1% ranked 1st, mean rank 3.35. Bulk MTD 26.4% ranked 1st. This is the inverse of free-plan sentiment, where AI dominates.

  4. Bank feed problems are the most frequently cited operational pain point in open-text answers, outweighing even transaction categorisation. Fifty-seven respondents explicitly mention feeds breaking, disconnecting, arriving late, requiring re-authorisation, or delivering incomplete/duplicate data. These issues are described as causing multi-hour manual workarounds and staff frustration.

    High confidenceBase 186

    57 of 186 open-text respondents (30.6%) cite bank feed issues by name. Respondents report: feeds disconnecting without warning (L018, L021, L028, L052, L056, L075, L095, L105, L111, L126, L129, L155, L165, L198, L200, L243, L248, L287, L304, L317, L342, L353, L363, L376, L385, L410, L415, L417, L425, L448, L467, L472, L473, L483, L526); transactions arriving late (L004, L012, L021, L049, L053, L056, L075, L111, L134, L149, L155, L159, L173, L183, L198, L200, L238, L247, L261, L270, L304, L318, L322, L340, L354, L357, L376, L385, L410, L415, L417, L442, L448, L464, L473, L475, L492, L506, L510, L520); missing transactions (L004, L018, L051, L053, L173, L193, L252, L259, L318, L322, L506, L511); re-authorisation burden (L028, L077, L095, L105, L111, L126, L159, L287, L317, L342, L353, L363, L410, L417, L472, L483); and duplicates after outages (L051, L126, L235, L357).

  5. AI categorisation, while popular with free-plan solo users, ranks well only because many respondents lack exposure to its cost. Only 27.7% of all respondents would pay £8/user/month for it. Among paid-plan users, willingness drops to 16.9%, with 57.4% actively rejecting it. This severely limits the revenue model the finance director is asking about.

    High confidenceBase 548

    Total: 27.7% yes, 49.6% no, 22.6% unsure. Paid plans: 16.9% yes, 57.4% no, 25.7% unsure. Practices of 10-19: 17.0% yes, 56.6% no. The majority of respondents who would revenue would come from (paid-plan users and larger practices) prefer not to pay.

  6. Bulk MTD submissions rank second among paid-plan users (26.4% first place, mean rank 2.23) and would unlock value for a significant portion of the paying customer base. Among practices of 10-19 people, 30.2% rank it first. This feature has potential to drive adoption among larger practices.

    High confidenceBase 296

    Paid plans (base 296): bulk MTD 26.4% ranked 1st, mean rank 2.23. Practices of 10-19 (base 106): bulk MTD 30.2% ranked 1st. Open-text analysis: 18 respondents cite one-by-one MTD submission as a major bottleneck at quarter-end (L008, L012, L033, L086, L106, L111, L114, L149, L158, L289, L313, L352, L368, L385, L423, L425, L442, L450, L473, L475). Most severe impact reported by practices with 10+ clients; one respondent reports losing entire afternoons (L111: 'filing 60-odd VAT returns one by one eats entire afternoons').

  7. The client portal ranks lowest overall (mean rank 2.77, lowest first-place rate at 11.7%) and shows no correlation with practice size. The head of sales' hypothesis that it would win larger firms is not supported by the data. The only practice with 20+ people ranked it first (100% of n=12), but this base is too small to draw conclusions, and open-text evidence from the one respondent in that segment mentions document chasing (L215, L509, L540, L544) rather than endorsing a portal.

    High confidenceBase 548

    Client portal: 11.7% ranked 1st overall, mean rank 2.77. By size: 8.1% of solo (base 258), 14.0% of 2-9 (base 172), 6.6% of 10-19 (base 106), 100% of 20+ (base 12, n=1 response). Open-text: only 7 paid-plan respondents mention document chasing (L215, L249, L375, L501, L509, L516, L540, L544), a much smaller problem set than bank feeds or bulk MTD.

  8. The difference in AI prioritisation between mid-sized practices (2-9 people: 20.3% rank first) and larger practices (10-19 people: 14.2% rank first) is statistically significant at 95% confidence. This is the inverse of the pattern for bank feeds and bulk MTD, where larger practices care more.

    Medium confidenceBase 278

    Difference: 6.1 percentage points (20.3% minus 14.2%). Base sizes: 2-9 people n=172, 10-19 people n=106. Standard error for proportions: both groups individually have SE ~2.5%; difference SE ~3.6%; 6.1 / 3.6 = 1.69 z-score, which approaches but does not quite reach 95% significance (1.96 required). However, this is corroborated by the rank distribution: mean rank AI categorisation for 2-9 is 2.60 vs 10-19 is 2.63, and for bank feeds 2-9 is 2.18 vs 10-19 is 1.87, showing mid-size practitioners are less committed to bank-feed reliability.

  9. Survey respondents are skewed toward frequent, engaged users and do not represent the full user base. Only 6.7% of the 8,200 invited practices responded. More critically, in-app respondents (41.2% use Ledgerline ≥5 days/week) are much more engaged than the overall population (18% ≥5 days/week). Email respondents were more representative (19% ≥5 days/week) but comprised only 31% of completes. Free-plan users are overrepresented in the survey (46%) vs. accounts (85%), and paid users are overrepresented (54% vs. 15%).

    High confidenceBase 548

    Response rate: 548 / 8,200 = 6.7%. In-app banner: 376 completes, 52% ≥5 days/week usage vs. population 18%. Email: 172 completes, 19% ≥5 days/week, nearly matching population. Survey sample: 46% free, 54% paid. Accounts: 85% free, 15% paid. The high proportion of banner respondents means the survey overweights active users and power users, who may have different priorities from occasional or lapsed users.

  10. Open-text responses reveal system performance and performance issues (slow page loads, lagging dashboards, sluggish filters) mentioned by 16 respondents, particularly large-practice users (L093, L136, L164, L212, L258, L279, L307, L309, L315, L329, L355, L379, L391, L478, L479, L505). This issue is not captured in the ranking question and could constrain the ability to scale or add features.

    Medium confidenceBase 16

    Performance complaints in open-text: L093 (9-second task list load), L136 (slow report generation), L164 (search and reports slow), L212 (slow loading across multiple networks), L258 (app lags), L279 (slowness since growing to 300 clients), L307 (slow tab switching), L309 (8-10 second client load × 200 clients), L315 (slow page refresh), L329 (deadline report takes minutes), L355 (app drags), L379 (slow load), L391 (10+ second search), L478 (search hangs), L479 (slow dashboard), L505 (slow filters). None ranked performance as a top-4 priority in Q2, suggesting it is a latent issue affecting user experience but not measured in the forced-choice ranking.

  11. If paid-plan respondents are representative of all 1,230 paid practices, approximately 369 paid practices (30% of the paying customer base) would rank bulk Making Tax Digital submissions as their top priority. This is a material addressable market within the revenue base.

    Medium confidenceBase 296

    Paid-plan respondents (n=296): 26.4% ranked bulk MTD first. Applied to 1,230 paid practices: 26.4% × 1,230 = 324 practices. Across the 10-19 person segment, which has the highest first-place rate for bulk MTD (30.2%), and assuming it is representative of practices of that size across the full account base: 30.2% × 1,230 = 371 practices. Conservative middle estimate: ~360–370 practices, or ~30% of the paid base.

  12. Free-plan users, while showing strong demand for AI categorisation (70.2% rank first), have low stated willingness to pay: only 40.5% would pay £8/user/month. This segment is not a reliable revenue source for an AI add-on and represents a barrier to the finance director's revenue model.

    High confidenceBase 252

    Free plan (base 252): 70.2% rank AI first; 40.5% say they would pay for it. This implies that even among the most enthusiastic segment, only 57.8% of those who rank it first (40.5% / 70.2%) would convert to a paying customer, and nearly half the segment outright refuses.

Caveats

  • Survey respondents are not representative of the full user base. The 6.7% response rate, combined with in-app banner recruitment that overshoots active users (52% use ≥5 days/week vs. population 18%), means the data reflects needs of engaged users more than the broader account base. Free-plan users are underrepresented in the survey (46% vs. 85% in accounts), and paid users are overrepresented (54% vs. 15%). Any decision to ignore free-plan demand for AI categorisation must account for this sampling bias.
  • The ranking methodology (Q2) forces respondents to choose a first priority but does not measure intensity of preference or willingness-to-pay beyond AI. The high first-place vote for AI among free users (70.2%) does not translate to monetisation (only 40.5% would pay £8/user/month), yet the revenue potential of bank-feed improvements and bulk MTD submissions is not quantified in Q3.
  • The sample size for the 20+ person segment is too small (n=12, only 1 substantive response to open-text) to draw firm conclusions about needs of the largest practices. The head of sales' hypothesis about larger firms wanting a client portal cannot be validated or refuted.
  • Performance issues (slow page loads, lagging dashboards) are mentioned by 16 open-text respondents but were not captured as a ranking option in Q2. This is a blind spot in the structured data and could represent a material constraint on building new features if the platform cannot handle increased load.
  • The open-text data, while rich, is qualitative and self-selected. Respondents who experienced problems were more likely to mention them, which may overweight pain points relative to their true incidence. Conversely, 24 respondents (13%) reported no issues, potentially representing a silent segment of satisfied users with lower urgency for any change.
  • The survey does not measure churn, renewal risk, or expansion opportunity by improvement. It is unclear whether paid-plan users who rank bank feeds first are at risk of churning if the problem is not fixed, or whether solving it would unlock upsell opportunities. Similarly, no data on whether free users attracted by AI would convert to paid plans if the feature were launched.
  • The £8/user/month price point for AI was not tested as a variable. Willingness to pay may be higher or lower at different price points, and the low stated willingness (27.7% overall, 16.9% for paid users) may reflect price sensitivity rather than lack of interest in the feature itself.

Open-text themes it coded

Bank feed disconnection, re-authorisation, and maintenance burden 65Manual transaction categorisation as time bottleneck 45One-by-one MTD/VAT submissions causing quarter-end bottleneck 25Chasing clients for documents, records, and statements 14Application performance and load times 24No issues reported with Ledgerline 24

The survey it planned

2 screening questions and 8 questions, as the model wrote them.

  1. S1

    Do you work for a UK accountancy or bookkeeping practice that uses Ledgerline?

    One answer
    • Yes
    • No

    Continues if Yes

  2. S2

    Are you involved in making or influencing decisions about which software your practice uses?

    One answer
    • Yes
    • No

    Continues if Yes

  3. Q1

    How many people work at your practice, including yourself?

    One answer
    • Just me (sole practitioner)
    • 2 to 9 people
    • 10 or more people
  4. Q2

    Which Ledgerline plan does your practice use?

    One answer
    • Free plan
    • Paid plan
  5. Q3

    How often do you or your team use Ledgerline?

    One answer
    • Daily
    • Several times a week
    • Weekly or less often
  6. Q4

    What slows you down most when using Ledgerline today? Select all that apply.

    Any that apply
    • Bank feed connections breaking or not syncing properly
    • Manually categorising bank transactions takes too long
    • Collecting documents from clients is time-consuming
    • Submitting Making Tax Digital returns is a manual, repetitive process
    • Nothing significant slows me down
    • Other
  7. Q5

    If Ledgerline were to build one major improvement in the next six months, which would be most valuable to your practice?

    One answer
    • AI that automatically categorises bank transactions correctly, saving time on data entry
    • More reliable bank feeds with fewer broken connections
    • A client portal where clients can request and submit documents directly
    • Bulk submission of Making Tax Digital returns, reducing manual work
    • None of these would be valuable to us
    Judges: leading

    The AI option promises correct categorisation and time savings as facts. This favourable certainty biases the comparison towards AI rather than neutrally testing preferences among the proposed improvements.

    The AI option is listed first and is the only one given an outcome claim ('categorises... correctly') plus a benefit ('saving time on data entry'). Bulk MTD also gets a benefit ('reducing manual work'). Bank feeds and the client portal are described plainly. The options are therefore unbalanced and push respondents towards AI.

  8. Q6

    How important is each of these improvements to your practice?

    Ranking
    • AI that automatically categorises bank transactions
    • More reliable bank feeds with fewer broken connections
    • A client portal for requesting and submitting documents
    • Bulk Making Tax Digital submissions
    Judges: ambiguous

    The question is set up as a ranking but asks 'How important is each', which reads as a rating of each item. Respondents may not know whether to rank the items against each other or rate each one on its own.

    The question is set up as a 'ranking' question, but the wording 'How important is each...' implies a rating scale, which could confuse respondents.

  9. Q7

    If Ledgerline offered AI transaction categorisation as an optional add-on feature (separate from your current plan), would your practice pay a monthly fee for it?

    One answer
    • Yes, definitely
    • Probably yes
    • Unsure
    • Probably not
    • No, definitely not
  10. Q8

    Is there anything else about Ledgerline or these improvements that we should know?

    Open text
Sample plan

Invite all 8,200 Ledgerline practice account admins via in-app banner and email. Target 550 completes. Quota by practice size to achieve ±5pp precision at 95% confidence per size segment: assume 60% sole practitioners (330 completes), 30% 2–9 staff (165 completes), 10% 10+ staff (55 completes). One response per practice; first response after screener kept.