Accounting, Tax & Bookkeeping// diagnostic

The outstanding-items board says complete while the preparer waits

In short

A tracker reads complete while work is blocked because one received flag answers only "did a file arrive", not "is it legible, does it cover the whole period, has anyone accepted it". Confirm with a re-open sample of completed items against the preparer's real blockers, then attach an acceptance test to each item type.

Key takeaways

  • The board is not wrong, it is narrow: one flag answers arrival and is read as readiness.
  • Confirm it with a re-open sample of 20 completed items before changing anything — the split between causes decides the fix.
  • Partial period coverage is the quietest cause: 11 monthly statements against a 12-month year reads as complete on every board that counts files rather than periods.
  • A superseded upload silently overwrites a corrected one when the item stores a single latest file with no version history.
  • Marking items complete to clear a personal queue is a metric problem, not a discipline problem: measure open items by engagement, not by assignee.
  • The fix is an acceptance test per item type, machine-checked where possible, not another status value in the same dropdown.

The board is not lying. It is answering a narrower question than the one the preparer is asking. Complete on most trackers means a file arrived against a request line; the preparer needs to know that the file is readable, covers the whole period, is the current version, and that somebody with the standing to say so has accepted it. One checkbox is standing in for four questions, and it answers the easiest of them.

The visible symptom is a trust failure rather than a data failure. Staff stop opening the board and go back to messaging the client manager directly, which is more expensive than the board ever saved. Once that has happened, adding a status column will not bring anyone back — the board has to be demonstrably right for a period before it gets read again.

Before changing anything: re-open 20 completed items

This is a 90-minute exercise and it decides which of the five causes below you actually have. Do not skip to the fix; the causes have different owners, and rebuilding for the wrong one is how a tracker gets a second status field nobody trusts either.

  1. Pick 3 engagements that are currently blocked and list the preparer's real blockers in their own words, before looking at the board.
  2. Pull every request item on those engagements marked complete — aim for about 20 — and open the evidence attached to each one.
  3. Score each against 4 tests: does it open, is it legible, does its coverage match the requested period, and is it the latest version the client sent.
  4. Record which test failed, not just that one did. The distribution across those four is the diagnosis.
  5. Compare the failures against the preparer's list from step 1. Blockers with no matching failed item are a different problem — usually a record that was never requested at all.

Cause 1: arrival and usability share one checkbox

The dominant cause. An item flips to complete when a file is attached, and nothing downstream ever revisits that decision. A password-protected statement, a photograph of a laptop screen, a scan with the last page cut off — all of them are received, and all of them block a preparer. The tracker records the client's action and the preparer needs the firm's judgement, which is a later and different event.

This is also where the machine-checkable tests live. A bank statement whose closing balance does not equal opening plus movements has failed its own arithmetic, which is the check behind the extracted statement does not foot; an invoice whose lines do not reconcile to its total fails the same way, as the lines do not sum to the invoice total sets out. Neither needs a human to detect. Any item type with an internal consistency rule should never be able to reach accepted without passing it.

Cause 2: eleven months counted as a year

The quietest one, because everything about it looks right. The client uploads 11 monthly statements, the item says "bank statements", a file is attached, done. Nothing in the item knows that the requested period was 12 months. The same shape appears with a statement that covers 1 January to 30 November for a year ending 31 December, or a payroll export that stops at the last completed quarter.

The fix is to store the requested coverage as a date range on the item and the delivered coverage as a date range on the evidence, then compare them. That comparison is free once both exist and impossible before. It is also the mechanism that lets next period's list be derived from what this one consumed, which is the argument in generating the request list from the last close.

Cause 3: the corrected file was overwritten by the wrong one

A client sends a corrected trial balance on Tuesday and re-forwards the original on Thursday because they lost the thread. If the item holds one latest file, Thursday wins. The board still reads complete, the preparer works from a superseded document, and the discrepancy surfaces at review, weeks later, as an unexplained difference.

Two changes prevent it. Keep every file received against an item, with its received timestamp and its source, so latest is a query rather than an overwrite. And treat re-receipt as a state transition that reopens the item for a second acceptance, rather than as a silent update — the same discipline that keeps records arriving on a reply thread from vanishing, described in emailed records never reach the engagement folder.

Cause 4: staff are closing their queue, not the engagement

Where a junior is measured on open items assigned to them, marking complete is the cheapest way to reduce the number. This is almost never a discipline problem and almost always a measurement one. It shows up in the sample as a concentration: most of the failed items were marked by the same one or two people, often in the same few minutes.

  • Measure open items per engagement, not per assignee. The engagement is what the client experiences and the partner is asked about.
  • Separate the two actions. Received is a fact anyone may record; accepted is a judgement, restricted to whoever will work the file.
  • Make reopening cheap and unblamed. If reopening an item is embarrassing, nobody reopens one, and every wrong close becomes permanent.

Cause 5: one upload satisfied an item that needed four files

"All loan statements" is a single line with an unknown number of documents behind it. The first attachment closes it. The same happens with "credit card statements" at a client with 4 cards, or "stock counts" across 3 locations. The item was written as a category when it should have been written as an artefact, and no state model fixes a request line whose completion is undefined.

The rule is one item, one artefact, one entity, one period — which usually means the list gets longer and closes faster. Where the count genuinely is not known up front, store an expected count that the client confirms, and treat the item as partially received until the count is met. The same discipline is what keeps an incoming handover honest in running a client file handover without silent losses.

Fix it as acceptance criteria, not as another dropdown value

The instinct after a bad board is to add statuses. That produces a longer dropdown and the same problem, because the defect is not the vocabulary — it is that nothing defines what would make an item done. Attach an acceptance test to each item type and let the state follow from the test. The transitions themselves are specified in the states a requested document passes through; what follows is what the tests look like.

Item typeAcceptance testRun by
Bank or card statementOpens without a password, period matches the request, closing balance equals opening plus movementsMachine
Purchase or sales invoiceLines sum to the stated total including tax, supplier identifiable, date inside the periodMachine
Payroll registerCovers every pay date in the period, gross to net reconciles per employeeMachine, exceptions to a human
Loan or lease scheduleOne document per facility, opening balance ties to the prior period closeHuman, count checked by machine
Stock count sheetsOne per location, signed, dated inside the count windowHuman
Acceptance tests by item type, and who or what can run them

Machine-checkable rows are the ones worth building first, because they are the ones failing silently today. A test that runs at upload also gives the client a chance to fix the problem while the document is still in front of them, which is the entire argument for pulling capture forward rather than waiting for review.

A status nobody can fail is not a status. It is a receipt for an upload, printed in the language of progress.

What separating the states will not fix

It will not make a client answer. Items that are genuinely outstanding stay outstanding, and the board will look worse for a period once the false completes are removed — which is the point, and worth telling the partner before you ship it. It also will not fix records that never made it into the request list, which is a generation problem, or records arriving by channels the board cannot see, which is an intake problem addressed in moving email-only clients onto a records portal.

It will not fix chasing either. A correct board tells you what to chase; how often and by whom is a separate design, covered across the client intake, chasing and portals topic. Building the acceptance tests and the state model into an existing practice stack is the kind of scope we take on as MVP and product builds for accounting and tax firms.

Frequently asked questions

Short answers to the follow-ups this page tends to raise.

How do we tell whether the tracker is wrong or the staff are marking things wrongly?

Look at the concentration in a re-open sample: a state-model fault spreads failures evenly across assignees, a behaviour fault clusters them on one or two people and often on one afternoon. Both can be true at once, but the ratio tells you where to spend first. If failures are spread and roughly 1 in 5 items fails, no amount of retraining will hold, because the system permits the wrong answer.

Should received and accepted be two fields or two states of one field?

Two states of one field, with timestamps and an actor on each transition. Two independent booleans permit combinations that are nonsense — accepted but never received — and someone will produce one within a month. A single state with a defined transition path also lets you measure the gap between arrival and acceptance, which is usually the real queue in a busy period.

Does an AI extraction step remove this problem?

No, but it changes where the answer comes from. Extraction can run the machine-checkable acceptance tests — period coverage, balance continuity, line arithmetic — and score confidence per field, which turns a subjective read into an automatic pass or fail. It does not decide whether a stock count sheet is signed by the right person. Route those to a human and keep the distinction visible in the state.

How long should it take to trust the board again?

One full cycle of whatever the firm's period is, with the false completes visibly removed. Trust here is empirical: staff resume using a board after they have opened it a few times and found it matched reality. Announcing the fix does nothing; showing a quarter-end where the board predicted the blockers correctly does most of it.

  • client intake
  • request tracking
  • engagement workflow
  • state machines
// shipped work

The work behind this page

Builds from our portfolio that this page draws on.

Read next

Working on something in this space?

Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.

Start the conversation