A no-heat call got booked as a routine tune-up
In short
When an intake agent books an emergency as routine, the cause is almost never model quality. It is that the tier policy was never written down in a form the agent could apply, or that nothing forces escalation when the agent is unsure. Joining a sample of booked jobs to their recordings and scoring them against your own tier table tells you which of the two you have.
Key takeaways
- Measure before tuning. Score 200 booked jobs against your own tier table, split by trade and by hour.
- Under-tiering and over-tiering have different causes. Fixing one blindly usually inflates the other.
- If two human scorers disagree on more than 1 in 10 calls, the policy is the defect, not the agent.
- An uncertain agent must escalate, not guess. Confidence has to gate the booking, not decorate the log.
- Question order matters: an agent that offers a slot before it establishes severity has already decided.
The observable is specific: a job on the board at routine priority, scheduled 6 days out, whose recording has a caller saying the heating has been off since yesterday and there is a baby in the house. Nothing errored. The booking is well formed, the customer was polite, and the agent's log shows a completed call.
The reflex is to blame the model and start rewriting the prompt. That is almost always the wrong first move, because two much duller causes account for most of these: the tier policy does not exist in a form anything can apply, and nothing forces the agent to escalate when it is unsure. Both are measurable in an afternoon.
Confirm the mis-tiering before you touch the agent
One bad booking is an anecdote, and a dispatcher who was already sceptical will supply three more. What you need is a rate, split by the two dimensions that actually move it.
- Pull every job booked by the agent over 14 days. If volume is high, sample 200 at random rather than taking the most recent 200, which over-weights whatever changed last week.
- Join each booking to its recording or transcript. Any booking you cannot join is its own finding: an agent whose decisions are unauditable cannot be improved, only replaced.
- Have 2 people who know the business score each call against the operator's own tier table, blind to the tier the agent chose.
- Where the 2 scorers disagree, keep the call in a separate pile. That pile is the single most useful artefact of the exercise.
- Compute under-tiering and over-tiering rates separately, then split both by trade and by hour of day, in 3 buckets: business hours, evening, and overnight.
- Read 10 mis-tiered transcripts end to end before drawing conclusions. The pattern is usually visible in the wording of the questions, not in the statistics.
What the numbers point at
| Measure | Investigate above | What it usually means |
|---|---|---|
| Under-tiered bookings | 5% of calls | Severity questions are missing or asked too late |
| Over-tiered bookings | 10% of calls | Keyword triggers rather than conditions; expensive but visible |
| Scorer disagreement | 10% of calls | The policy itself is ambiguous. Fix that first |
| Overnight versus daytime gap | 2x the daytime rate | A different path runs after hours, or capacity is empty |
| Calls with no severity question asked | 5% of calls | Question order books before it triages |
| Emergencies booked with no free slot | any | Capacity pressure is silently rewriting the tier |
The overnight split matters more than teams expect, because after-hours calls often run through a different path, a different prompt or a different fallback entirely — the same seam that produces overnight bookings missing from the morning board.
Five causes, ranked by how often they are the one
| Cause | Signature in the audit | Fix | Owner |
|---|---|---|---|
| Tier policy never written down | High scorer disagreement; errors spread evenly across hours | Write the conditions as yes-or-no questions with a response commitment | Operations, not engineering |
| Caller language hides severity | Errors cluster on vague openings; scorers agree once they hear the call | Add 2 explicit severity questions the agent must ask regardless of wording | Whoever owns the script |
| No confidence threshold | Mis-tiered calls read as fluent and decisive | Force escalation below a set confidence, and log the reason | Engineering |
| Capacity pressure | Errors spike when the board is full; emergencies land in the only slot offered | Separate tier from slot; make an unschedulable emergency an alert | Dispatch |
| Question order | Slot offered before any severity question in the transcript | Establish tier first, then offer only slots that tier permits | Whoever owns the flow |
When the policy was never actually written
This is the most common finding and the least popular one, because it is not a software fix. If 2 experienced people listening to the same call assign different tiers, no agent — human or otherwise — can be consistent, and every prompt change simply moves the errors around. The disagreement pile from the audit is the agenda: work through it with whoever sets policy and turn each contested call into a condition that can be asked about.
Write the conditions in the caller's language, not the trade's. 'Is the heating producing no warm air at all, or is it running but not keeping up?' is answerable by a homeowner in a cold kitchen. 'Is the system in hard lockout?' is not. The definitions themselves belong on their own page — see emergency, urgent, same-day and routine — but the version that ships is the operator's, dated and owned.
When the agent is confident and wrong
If the policy is clean and scorers agree, the failure is that nothing stops the agent from proceeding when it should not. Fluency is not calibration: an intake agent will state a tier as confidently on a call it half-understood as on one it read perfectly. The containment is a threshold with an action attached, and the action has to be a real branch in the flow, not a note in the log.
Two adjacent faults distort this measurement, so rule them out before concluding the threshold is wrong. If the caller is not matched to their history, an agent cannot know this is the third visit for the same fault — the mechanism is in why every repeat caller becomes a new customer record. And if the audit is drawn from bookings alone, it silently excludes the calls that never became bookings, which is the reconciliation problem behind an ad account counting more leads than the board.
Question order decides the tier before anyone notices
Read the mis-tiered transcripts and one pattern recurs: the agent offered a time before it established what was wrong. Once a slot is on the table the conversation is about the slot, the caller accepts what is offered, and the tier is backfilled to match. Establish the condition first, then offer only the windows that tier permits.
There is a real cost to this, and it should be measured rather than assumed. Every question before the first useful answer is time a caller may not give you, and abandonment in the opening seconds is its own failure — see callers hanging up before the greeting finishes. Keep the severity path to 2 or 3 questions and let the rest of the record be completed after the tier is fixed.
The decision tree
- Do 2 scorers disagree on more than 1 in 10 calls? Fix the policy. Nothing downstream can be consistent until they agree.
- Do they agree, and the agent still under-tiers? Check whether the severity questions were asked at all. If they were not, the fault is question order, not judgement.
- Were the questions asked and answered, and the tier still wrong? Add the confidence gate and the escalation branch, then re-run the same audit in 14 days.
- Do errors spike only when the board is full? The tier is being rewritten by capacity. Separate tier from slot and alert on emergencies with nowhere to go.
- Do errors spike only overnight? Compare the after-hours path against the daytime one; you are probably auditing a system that is not the one taking those calls.
Re-run the audit on the same protocol after each change, and keep the transcripts. The point is not to reach zero mis-tiered bookings, which is unattainable when the underlying information is what a stranger says on a phone. It is to make the residual errors visible, cheap and owned. That measure-then-contain loop is how we scope intake work in AI agents and automation, and the rest of this cluster sits under intake and call capture, within our field service and trades work.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
How do I know whether my intake agent is booking the wrong priority?
Join a random sample of about 200 bookings from the last 14 days to their recordings, and have 2 people who know the business score each call against your own tier table without seeing the tier the agent chose. Report under-tiering and over-tiering separately, split by trade and by hour. A rate you have measured is arguable; a dispatcher's memory of last Tuesday is not.
Should I just add emergency keywords to the prompt?
No, and it usually makes things worse. Keyword triggers over-fire in season — 'no heat' shows up in ordinary maintenance calls all January — while the genuinely urgent caller often describes the problem without using any trigger word. Gate on whether the severity questions were answered and whether the decision cleared a confidence threshold instead.
What should the agent do when it is not sure how urgent a call is?
Take the booking details, avoid committing to a time, tell the caller plainly that a dispatcher will confirm the timing shortly, and raise an alert naming the unresolved condition. That converts an invisible wrong answer into a 2-minute human decision, which is the cheapest outcome available once the agent is genuinely uncertain.
Is over-tiering as bad as under-tiering?
It is expensive but far less dangerous, and it is visible: over-tiering shows up as overtime, disrupted routes and technicians arriving at non-emergencies. Under-tiering shows up as a customer without heat for 6 days and a review nobody can undo. Tune the thresholds asymmetrically, and accept a higher over-tiering rate than under-tiering rate.
- triage
- voice agents
- intake
- diagnostics
The work behind this page
Builds from our portfolio that this page draws on.
FieldRoute
An AI field-service platform that auto-dispatches the best-matched technician, optimizes routes, and tracks first-time-fix against every SLA.
OperationsHaulBoard
An AI freight load board that matches every open load to the best-fit carrier, prices each lane on live spot-rate data, and tracks broker margin on every move.
LogisticsRead next
- Emergency, urgent, same-day, routine: four tiers and who defines themThere is no industry-wide definition of an emergency call. The operator writes one, and until those conditions exist as data no intake agent applies them consistently.definition
- The job type is the schema: what a booking code has to carryDispatch, duration, skill matching and pricing all read the job type. Treating it as a dropdown label rather than a schema is why the board looks full and the day falls apart.definition
- A home warranty dispatch is not a customer callA warranty dispatch arrives with an authorisation number, a covered scope someone else defined, and a payer who is not in the house. Booking it like a retail call produces uncloseable jobs.definition
- The calls were answered overnight and the board is empty at 7amThe overnight calls were answered and the intent captured, and the jobs still vanished in the handoff. Four counts, five ranked causes, and the report that lands before the board opens.diagnostic
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation