Insurance & Claims// diagnostic

Fraud indicators fire on a third of the book and nobody reads them any more

In short

When fraud indicators fire on a third of the book, the first move is not retuning thresholds — it is producing the join almost no claims operation has: per-indicator firing rate against referral acceptance against investigation outcome. Without it, tuning is an argument between opinions. With it, each indicator resolves to retire, re-threshold or combine.

Key takeaways

  • The confirming check is a join: indicator fired, referral accepted, investigation outcome — most operations have never produced it.
  • An indicator that fires on a third of the book and appears in a third of accepted referrals carries no information at all.
  • Correlated indicators counted separately turn a single circumstance into 4 points of apparent evidence.
  • Acceptance rate alone is a poor target because it moves with investigator capacity, not with claim behaviour.
  • Every indicator resolves to retire, re-threshold or combine, and the choice follows from lift rather than from intuition.
  • The output is a referral input for a trained investigator and never a determination that a claim is fraudulent.

The visible problem is that investigators have stopped reading the alerts. The actual problem is that nobody in the operation can say what a single indicator buys: how often it fires, how often the referrals it drove were accepted, and how often those investigations substantiated anything. Until that join exists, every threshold change is an argument between confident people with no evidence.

So the first finding of this diagnosis is usually the measurement gap itself. Fraud indicator sets are typically inherited — from a vendor default, a predecessor's spreadsheet, a regulator's published list, or a model built years ago on a book that no longer looks like this one — and are then never evaluated per indicator, only in aggregate.

The join nobody has produced: indicator, referral, outcome

The evidence base is one table, at indicator granularity, over a period long enough for investigations to have closed. Building it is unglamorous and is nearly always the highest-value week in the project.

  1. Log every indicator evaluation, not just the ones that fired. A rule that fires on 31% of files is only interpretable against the 100% it was evaluated on, and most systems record only the hits.
  2. Attach the referral decision. For each file, whether a referral was made, whether the investigation unit accepted it, and the stated reason where one was captured.
  3. Attach the closing outcome. Substantiated, unsubstantiated, or closed without a conclusion — and keep the third category separate, because merging it into either of the others biases every number that follows.
  4. Wait out the lag. Investigations close weeks or months after referral, so a cohort younger than the median investigation duration will look worse than it is. Cut the cohort by referral date, not by report date.
  5. Compute lift, not counts. For each indicator: firing rate across all files, presence rate among accepted referrals, presence rate among substantiated outcomes. An indicator whose 3 rates are the same is telling you nothing.
IndicatorFires onPresent in accepted referralsReading
Policy under 12 months old34%35%No lift; it describes the book, not the claim
Loss reported 30+ days after the date of loss11%24%Real lift; keep, and check the date quality behind it
Prior claim within 24 months19%21%Marginal; likely correlated with something already counted
Bank details changed within 7 days of payment1%12%Rare and sharp; weight it far above the others
The shape of the finished table, on 4 illustrative indicators — the arithmetic, not measured data

Row 1 is the pattern that makes a whole indicator set unreadable. If a third of the portfolio is under 12 months old, an indicator for policy age fires on a third of the book by construction, and it will appear in roughly a third of everything else too. It has cost the operation attention and returned nothing.

5 reasons a third of the book lights up

CauseConfirming checkFix
Indicators that are population characteristics, not anomaliesFiring rate matches the base rate of that attribute across the bookRetire, or re-express as a deviation from the segment's own norm
Correlated indicators counted as independent evidencePairwise co-occurrence among fired indicators is high on most flagged filesGroup into families and score the family once
Thresholds set against a different bookFiring rates differ sharply by region, line or channel from what the rule assumedRe-fit per segment, or make the threshold a segment-level parameter
No decay on historical indicatorsThe same file keeps firing the same prior-claim indicator years laterTime-bound every backward-looking indicator and let it expire
Scoring that sums instead of weightingEvery indicator contributes the same point, so volume beats severityWeight by measured lift, and cap the contribution of any one family
Ranked causes of indicator flood, the check that confirms each, and the fix

One circumstance, counted 4 times

Correlation is the cause that survives longest, because each indicator looks defensible on its own. A loss that occurs at 2am on an unlit road will often also have no independent witness, no contemporaneous police attendance, and a delayed report. Those are 4 indicators and 1 circumstance. A summing score reads it as 4 pieces of evidence, and the file arrives at the investigation unit looking far stronger than it is.

The fix is structural rather than statistical. Group indicators into families — circumstance of loss, claimant history, documentation quality, payment behaviour, provider or repairer pattern — and let a family contribute once, at the weight of its strongest member. The score stops being a count of how many boxes a file ticks and starts being a count of how many independent things are odd about it.

Retire, re-threshold or combine: the decision per indicator

Once the join exists, each indicator lands in exactly one of 4 places. The point of writing it as a decision rather than a review is that it can be re-run every quarter without another workshop.

  1. Fires often, no lift in accepted referrals. Retire it. Keep it logged as a field if anyone wants it back, but take it out of the score — this is where most of the noise lives.
  2. Real lift, but fires far too often to action. Re-threshold, and prefer a segment-relative threshold over a global one. A 30-day reporting delay means something different on a commercial property book than on personal motor.
  3. Lift only in the company of another indicator. Combine. Make it a member of a family, or express the pair as a single composite condition, so it cannot contribute twice.
  4. Rare and sharply predictive. Keep and raise its weight. Rare indicators are what the whole exercise is trying to protect: they are drowned by the common ones long before an investigator gets to them.

An indicator earns its place by changing what somebody does. One that fires on a third of the book changes nothing except how quickly the alert panel gets closed.

What tuning will not fix

  • Bad dates underneath good rules. Reporting-delay and duration indicators are arithmetic on fields that are routinely wrong or absent, which is why the 3 dates every claim record needs has to be settled before any delay-based indicator is trusted.
  • A queue that nobody works to the bottom. If the investigation unit's capacity is the binding constraint, a better-ordered queue helps and a longer one does not — and the files that quietly rot are rarely the newest, as the files in the middle of the caseload shows.
  • Reserves that never move after a referral. An accepted referral changes the expected cost of the file, and if the reserve does not respond, the operation's own numbers stop reflecting what it knows — the trigger problem in reserves that get set once and never move.
  • Confusing this screen with the recovery screen. Fraud screening and recovery screening ask different questions on different clocks, and recovery has a hard statutory deadline the fraud screen does not, covered in finding recovery potential before the limitation runs.
  • A score presented instead of an action. Anything that hands a human a number to interpret rather than a next step gets ignored, which is the general lesson in AI in logistics operations and holds just as firmly in a claims unit.

The line the output must never cross

An indicator score is a prompt to look, produced for a trained investigator. It is not a finding that a claim is fraudulent, and it must never function as one. That is a design constraint with concrete consequences: the score should not on its own delay a payment, decline a claim, or attach a label to a policyholder record that follows them into renewal or underwriting.

  • Every score shows its constituent indicators and their weights, so a person can see what drove it and disagree with it in writing.
  • Referral, acceptance and outcome are recorded as separate events with timestamps and actors, because a decision nobody can reconstruct cannot be defended to a regulator, an ombudsman or a court.
  • Indicator sets and weights are versioned, so a file can be re-examined against the rules that were live on the day it was flagged rather than today's.

None of this requires a new model. It requires logging every evaluation, joining 3 systems that already hold the data, and writing the tuning rule down so it survives the person who wrote it — AI agents and automation work in the unfashionable sense. This page sits in claims handling, fraud flags and recovery, part of insurance and claims software.

Frequently asked questions

Short answers to the follow-ups this page tends to raise.

Why do fraud indicators produce so many false positives on claims?

Usually because several of the indicators describe the book rather than the claim. An indicator for a policy under 12 months old fires on whatever share of the portfolio is under 12 months old, regardless of behaviour, and contributes a point to every one of those files. Add correlated indicators counted separately and a summing score, and a third of the book crosses the threshold without any of it meaning anything.

What should be measured before tuning fraud indicator thresholds?

Three rates per indicator: how often it fires across all files evaluated, how often it is present among referrals the investigation unit accepted, and how often it is present among investigations that substantiated something. An indicator whose 3 rates are roughly equal has no lift and should be retired rather than re-thresholded.

How do correlated fraud indicators distort a claim score?

They convert one circumstance into several points of apparent evidence. A night-time loss with no witness, no police attendance and a late report is 4 indicators describing a single event. Grouping indicators into families and scoring each family once — at the weight of its strongest member — stops a score measuring how many boxes a file ticks instead of how many independent things are odd.

Can a fraud score be used to decline or delay a claim?

No. The score is a referral input for a trained investigator and carries no finding about the claim or the claimant. It should not by itself delay a payment, decline a claim, or attach a label that follows a policyholder into renewal. Keep the contributing indicators visible, log referral and outcome as separate events with actors and timestamps, and version the indicator set so a decision can be reconstructed later.

  • fraud indicators
  • SIU referral
  • alert fatigue
  • precision
// shipped work

The work behind this page

Builds from our portfolio that this page draws on.

Working on something in this space?

Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.

Start the conversation