False reject rate: what the line feels, not the accuracy on the slide
In short
False reject rate — overkill — is the share of conforming parts an inspection system rejects; its twin is the escape, a non-conforming part it passes. The floor experiences neither as a percentage. It experiences good parts in the reject bin and the minutes spent re-checking them, which is why the rate must be quoted in parts per shift before anyone signs for it.
Key takeaways
- False reject (overkill) and escape are the 2 errors, and they have different owners: production pays for one, the customer finds the other.
- Convert the rate before you accept it. At 1,200 parts per hour, 1% overkill is 96 good parts a shift and about 72 minutes of re-checking.
- Quote the rate per part presented, not per image. A cell taking 4 views of every part rejects at roughly 4 times the per-image rate.
- The re-inspection loop needs a named owner and a logged verdict, because that log is the only honest source of the true overkill number.
False reject rate, also called overkill, is the proportion of conforming parts an inspection system rejects. Its twin is the escape: a non-conforming part it passes. Both are error rates, but only one lands on the floor that shift, and that asymmetry is why the false reject rate decides whether operators keep using the cell.
Headline accuracy hides it, because defect rates are small. If 1 part in 500 is genuinely non-conforming, a system that simply passed everything still scores 99.8%. Accuracy on an imbalanced population is a statement about how rare defects are, not about how well the cell judges them.
Two errors, two owners, two discovery times
| Error | What happened | Who absorbs it | When it surfaces |
|---|---|---|---|
| False reject (overkill) | A conforming part was rejected | Production: yield, re-check labour, operator patience | Same shift, visibly |
| Escape (false accept) | A non-conforming part was passed | The customer, then warranty and containment | Weeks or months later |
| No-read | No confident call was possible at all | Whoever staffs the review station | Immediately, as a queue |
The two rates move against each other: tightening sensitivity to shrink escapes grows overkill, and loosening it does the reverse. That is why the setting is a governed choice rather than a tuning exercise — see the anomaly score is a ranking, the threshold a business decision.
Convert the percentage into a pile
A percentage is not reviewable by the people who live with it. Convert it against real throughput. The figures below assume 1,200 parts per hour on an 8-hour shift — 9,600 parts — and 45 seconds to retrieve, re-judge and dispose of each pulled part.
| False reject rate | Good parts pulled per shift | Re-check time at 45 s each |
|---|---|---|
| 0.1% | 10 | About 7 minutes |
| 0.5% | 48 | About 36 minutes |
| 1% | 96 | About 72 minutes |
| 2% | 192 | About 2.4 hours |
| 5% | 480 | About 6 hours — most of a person's shift |
Three things become arguable once the table exists: whether the re-check is staffed, whether the rate is acceptable per class or only on average, and whether anyone will notice the gap between 0.5% and 2% before the trial ends. Systems trained on conforming parts alone start high on overkill and improve as real rejects accumulate — the migration in when you have forty defect images, and when you have four thousand.
The re-inspection loop, and who staffs it
- The part leaves the line to a destination that is neither the scrap bin nor a tray on the floor, or the evidence is gone before anyone can use it.
- Someone judges it against the class list, with the golden and limit samples in reach — a named role on the shift plan, not whoever is free.
- The verdict is recorded against the part's image and score, agreed or overturned, with the class. This is the only record a true overkill rate can be computed from.
- The disposition is executed, and the queue depth is watched: a queue that grows all shift means the rate is above what the station can absorb.
If that loop is unstaffed the failure is quiet and predictable: pulled parts accumulate, someone starts returning them unchecked, and within a fortnight the cell is worked around rather than used. The station also needs somewhere to log a verdict in seconds — ordinary internal tools and ops work, nothing to do with the model.
A system is not abandoned in a meeting. It is abandoned by an operator who has learned that the reject bin is mostly good parts.
The number to agree before the build starts
Write the overkill budget in the units above and per class, not as one percentage: the acceptable pile for a class dispositioned scrap differs from one dispositioned use-as-is. Agree who staffs the re-check at that volume, and what happens if the rate lands at twice the budget in week 1. Trials that skip this run months on one line with still no decision, because nobody wrote down what good looked like in shift units.
The rest of the imaging, threshold and acceptance decisions sit under visual inspection and defect detection, and the on-premise vision work under manufacturing and industrial vision.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
What is an acceptable false reject rate for an inspection line?
There is no universal figure — it is whatever the re-check station can absorb while still beating the manual baseline. Convert candidate rates into parts per shift, price them in re-check minutes, and take the highest rate production will tolerate for the escape reduction it buys. A rate producing more pulled parts than the station can judge is unacceptable however good the percentage looks.
What is the difference between false reject rate and overkill rate?
They are the same measurement under two names. Statistics and quality practice say false positive or false reject; machine vision and semiconductor practice say overkill. The distinction worth policing is the denominator: overkill is usually quoted per part presented, a model's false positive rate per image or per region, and a multi-view cell makes those two numbers materially different.
How do you measure the real false reject rate after go-live?
By re-judging every rejected part against the class list and recording the verdict beside its image and score. The rate is overturned rejects divided by conforming parts presented, and it cannot be derived from the model's own confidence figures. That is why the loop must be staffed and logged from day 1: without it the only number available is the rejection rate, which conflates good parts with bad.
Does lowering the false reject rate always increase escapes?
Moving the threshold does — that trade is fixed by the model you have. Improving the images does not: better lighting, tighter fixturing or an extra view lowers overkill and escapes together, because they add information rather than reallocating it. Before conceding the trade, check whether the borderline rejects share an imaging cause; a cluster on one face or one shift usually means the picture is the problem.
- overkill
- inspection metrics
- line operations
- quality cost
The work behind this page
Builds from our portfolio that this page draws on.
Open Vision PPE Monitoring
Boundary surveillance, PPE compliance monitoring, and intrusion detection via real-time video analytics. Runs fully on-premise — no cloud required.
Safety & ComplianceFactory OS
Production planning and task management for a tier-1 apparel manufacturer — replacing Excel with automated milestone planning, SOP gate enforcement, and real-time visibility.
ManufacturingRead next
- The anomaly score is a ranking; the threshold is a business decisionThe score orders parts by how far they sit from normal. It carries no units, no probability and no severity — which is why the threshold belongs to quality, not to engineering.definition
- The golden sample is a decision record, not just a good partThe part in the drawer is the easy half. Its record — revision, attribute, approver, date, recheck trigger — is what an automated inspection inherits.definition
- Writing the defect class list your inspection will be graded againstEvery class needs a name, a definition, a severity, a disposition and a minimum size. Classes two inspectors cannot separate produce a confusion matrix nobody can act on.definition
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation