A bookkeeping pilot passed, then the team went back to the spreadsheet
In short
A bookkeeping pilot that passed and then died in production usually failed on placement, not training: the tool sits outside the window staff work in, reviewing its output costs more than doing the work, its exception queue has no owner, or it was trialled on clean records and met the real ones at go-live. Measure usage per preparer against eligible volume first.
Key takeaways
- Measure adoption as items processed divided by items eligible, per preparer, per week. "Nobody uses it" is usually two people using it on one client.
- A tool living in a second browser tab loses to the ledger the preparer already has open, however good its output is.
- If verifying a suggestion costs more than about 60% of the manual time, staff stop verifying and then stop using it.
- An unowned exception queue grows silently, and once its oldest item predates a filed period, nobody trusts the tool's completeness again.
- Pilots run on tidy clients and production is the shoebox. Trial the worst 10% of documents or the agreed pass mark measures nothing.
- The change that most often revives adoption is writing output back as a draft in the system of record, not a suggestion on another screen.
The pilot did not fail. The placement did. A tool that produced good output in a controlled trial and then quietly stopped being used at go-live was positioned beside the work rather than inside it, and the four causes are checkable in an afternoon. Training is the diagnosis firms reach for because it is the only lever they know how to pull, and it is the one that changes nothing.
The pattern is consistent: a partner asks how the new system is going, hears "we're still getting used to it", sees a dashboard confirming the tool is live, and does not look again until renewal. Meanwhile the preparers have rebuilt the old workbook.
Measure first: usage per preparer against eligible volume
Adoption is a ratio, not an impression. The numerator is items the tool processed. The denominator is items that should have flowed through it in the same period — bank lines on connected clients, supplier invoices received, documents uploaded. Compute it per preparer per week, never firm-wide, because the average hides the shape that identifies the cause.
- Fix the eligible set. Count items in the ledger or intake queue that met the tool's stated scope for the period — not total workload, only what it was bought to handle.
- Pull the processed set. Count items carrying a tool-created artefact: a suggestion accepted, a coded line, an extracted document, a posted entry.
- Split both by preparer and by week. Twelve weeks post go-live is enough; four is not, because the first two are training and the third is enthusiasm.
- Split the processed set by client. If more than half the volume sits on two or three files, you have a pocket of adoption, not adoption.
- Track the exception queue on the same weekly grid, recording the age of the oldest open item rather than only the count.
Four causes, ranked, with the signal that separates them
Run these in order. The first is the most common by a wide margin and the cheapest to test, and each has a signal you can read off the measurement you just built rather than off a staff survey.
| Cause | What the usage curve looks like | The confirming question | Direction of the fix |
|---|---|---|---|
| Wrong window | Drops steadily from week 3, evenly across staff | Which application is open on the preparer's second monitor all day, and is it this one? | Move the output into that application, or accept a second-screen tax you will keep paying |
| Review costs more than doing | Flat and low from day one; high accept-without-checking rate | Time both paths on the same twenty items with a stopwatch | Show the evidence beside the suggestion, or raise the auto-post threshold and stop asking |
| Unowned exception queue | Healthy for 4-6 weeks, then collapses; queue age climbs monotonically | Name the person who cleared the queue last Tuesday | One named owner, a daily clearing window, and an age alert |
| Pilot documents were not production documents | Falls fastest on the newest and messiest client files | What share of go-live exceptions come from clients that were not in the trial? | Re-trial on the worst documents and re-agree the pass mark |
Cause 1: the tool is not in the window the preparer lives in
A preparer spends the day inside one or two applications — the ledger, the practice management system, sometimes a workpaper file. Anything requiring them to leave those for a third screen competes not on its own quality but against the friction of the switch, and it loses that fight on every busy day. Busy days are the only ones that matter, because that is where the volume is.
The tell is a decay curve rather than a cliff, spread evenly across staff, collapsing first in the week before a filing deadline. If the tool's entire interface is its own dashboard, you have this cause until proven otherwise — the same logic behind buying extraction and building the review layer yourself. The part worth owning is the surface staff touch, because that is the part that has to sit where they already are.
Cause 2: verifying the suggestion costs more than doing the work
A preparer who knows a client codes a bank line in a few seconds. If checking the tool's proposed code means opening a source document in another tab, finding the reference, and coming back, verification can take four or five times as long as the original task. Nobody sustains that trade, and the failure is not that they stop using the tool — it is that they stop verifying first, accept a run of suggestions blind, get burned once at review, and abandon it in the same week.
- The threshold in practice. If verification costs more than roughly 60% of the manual time, the step is not a saving and staff will price that correctly before you do.
- The high-accept tell. An acceptance rate close to 100% with no correction history means suggestions are being waved through, not reviewed, and the review gate is decorative.
- The fix is evidence placement. The extracted value and the region of the document it came from belong on one screen, or the reviewer is doing the retrieval the tool was meant to remove.
- The other fix is fewer decisions. Auto-post above a confidence threshold and route only genuine ambiguity to a person — exceptions-first rather than review-everything.
Cause 3: the exception queue has no owner, so it poisons trust
Every extraction or matching system produces items it cannot resolve. That is correct behaviour. The defect is organisational: the queue is created with no named owner and no clearing window, so it accumulates until a preparer opens it and finds an item from a period already filed. From that moment the tool is nobody's system of record, and the spreadsheet returns as the thing that can be trusted to be complete.
This is the same defect as a status board that reports done while a person is still waiting, which is why the fix rhymes with the one described in why an outstanding-items board says complete when it is not: a single flag collapses states that need to stay separate. Track queue age, not queue depth. Depth looks fine right up to the point where it does not.
Cause 4: the pilot ran on clean records and production got the shoebox
Firms nominate their tidiest clients for a trial, for reasonable reasons — those clients are on the portal, their statements arrive as machine-readable files, their invoices are digital. Then go-live includes the client whose records are photographs of receipts taken at an angle, a statement scanned from a printout, and a payroll summary emailed as a screenshot. A pass mark set on the first population tells you nothing about the second.
If the worst tenth of your documents generates most of the exceptions, a trial that excluded them measured the wrong thing. The corrective is to re-run it deliberately on the hard set, which is the method behind running a vendor trial on your messiest client records — including agreeing the scoring sheet before anyone sees a result.
The decision tree, in order
- Is the ratio above 0.5 for anyone? If yes the tool works, so this is placement or assignment — go to step 2. If no, go to step 4.
- Is that usage concentrated on two or three client files? Find what is different about them — usually a document format or one preparer's habit — and decide whether it generalises.
- Did usage decay from about week three rather than never start? That is the wrong-window cause. Stop training and start moving the output.
- Time twenty items both ways with a stopwatch. If review is slower than roughly 60% of manual, fix evidence placement or raise the auto-post threshold first.
- Plot exception-queue age by week. If it climbs without a reset, name an owner and a daily clearing window today; nothing else holds until that is true.
- Compare exception sources against the trial population. Mostly from clients outside the trial means a scope problem, not an adoption problem.
One case sits outside the tree. Where two systems both claim the same step — a statement read twice by two products that disagree — abandonment is a symptom of the overlap rather than of either tool, and the resolution is the data-ownership seam in deciding between two tools that both read the same statements. The surrounding decisions sit in the build-versus-buy and vendor-risk topic.
The one change that most often brings it back
Stop asking staff to go and look at output. Write it into the system they are already in, as a draft they approve or amend in place — a draft ledger entry, a pre-coded line awaiting confirmation, a populated field on the practice record. The tool goes invisible and the work gets faster, the only order in which adoption survives a filing deadline.
Adoption is not a function of how good the output is. It is a function of how far the person has to travel to meet it.
That is usually a build rather than a setting, because it is the surface most vendors do not expose: a write-back layer and a review screen on top of a bought extraction engine. It is the shape of work we scope as an MVP and product build, inside the wider picture of software and AI systems for accounting and tax practices. Upstream of go-live, the sequencing argument in taking an AI proof of concept to production is worth reading before the pilot rather than after it.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
How long after go-live should we measure adoption?
Twelve weeks, measured weekly, starting from go-live rather than from the pilot. The first two weeks are training and the third is novelty, so a four-week read almost always looks healthy and tells you nothing. The decay pattern that identifies the wrong-window cause typically only becomes visible in weeks four to eight, and the exception-queue collapse usually lands somewhere between weeks four and six.
Is low adoption ever genuinely a training problem?
Occasionally, and it has a distinctive signature: high usage among staff who attended the sessions, near zero for everyone who joined afterwards, and no decay curve in either group. If usage decayed for people who were once using it well, training is not the cause — they demonstrably knew how.
Should we switch vendors if the pilot did not stick?
Not until you know which of the four causes you have, because three of them will reproduce with any vendor. Wrong window, unowned exception queue and an unrepresentative trial population are all properties of how the tool was placed and run, not of the tool. Only a genuine quality failure on your real documents — measured on the hard set, against a pass mark agreed in advance — justifies a switch.
What if staff say they like it but still do not use it?
Believe the usage ratio, not the survey. Staff are consistently polite about tools their firm has paid for, and they are also accurate about their own time under pressure. The gap between stated approval and actual use is almost always the review-cost cause: the tool is genuinely good, and verifying it is still slower than the thing they can already do without thinking.
Can we just mandate use of the tool?
You can, and it works only where the tool is not slower than the alternative. A mandate on a step that costs more time than it saves converts quiet abandonment into quiet compliance: suggestions accepted without review, which is worse, because the errors now carry an approval record.
- adoption
- pilots
- workflow design
- practice operations
The work behind this page
Builds from our portfolio that this page draws on.
Read next
- Gated endpoints: the ledger APIs you cannot call until your app is reviewedSandbox access proves nothing about production. Ledger platforms tier access, and the capabilities firms most want — payroll, bank feeds, filing — sit at the top tier.definition
- Professional clearance: the letter, and the records that have to moveThe clearance letter asks one question and takes 10 minutes. The data handover behind it decides whether the incoming firm can open a period, and it is where a change of accountant actually fails.definition
- The prepared-by-client list is a data structure, not a spreadsheetPBC stands for prepared by client. The useful definition is structural: a list of request items, each carrying nine fields, one of which is a state machine.definition
- What a client due-diligence record has to contain before work startsDue diligence is usually written about as a compliance narrative. The deliverable is a record: 10 fields, each with evidence attached, a refresh clock and a named approver.definition
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation