Rejects double the morning a new coil or resin lot goes on
In short
A jump in false rejects after a material batch change usually means the new lot looks different, not worse — a gloss, colour, coating or finish shift that moves the score distribution while parts stay good. Overlay reject rate on lot changeovers from goods receipt: a step at a changeover names the material. Then hand-inspect 30 rejects, because the case that matters is a lot that really is worse.
Key takeaways
- Overlay reject rate on goods-receipt changeovers first. A step aligned to a lot boundary names the variable.
- Hand-inspect 30 rejects before anything else. It separates a worse lot from a merely different-looking one.
- A monochrome camera can shift by tens of grey levels on a colour change the eye cannot see at all.
- Retraining on a bad lot teaches the model to accept it. Revalidate per lot instead of widening normal.
- None of this is possible unless the lot number reaches the image record at capture time.
The material did not necessarily get worse. It got different, and an inspection system trained on how the old lot looked is measuring departure from that appearance, not conformance to a drawing. A coil 5 gloss units brighter, a resin with a fractionally different pigment load, a plated layer 2 micrometres thicker — each of these moves the whole population of scores at once, and a fixed threshold that sat comfortably above the old distribution now cuts through the middle of the new one. Every part is affected simultaneously, which is why the pattern is a step from 1.2 percent to 5 percent at 06:00 rather than a drift over a week.
That is the likely explanation, not the certain one. The awkward possibility is that the lot really is defective and the system is doing exactly what it was bought for. Both hypotheses produce the same headline number, and the two checks below separate them in a single shift.
The overlay that turns 'it got worse' into a timestamp
- Pull hourly reject rate per part number for the last 60 days. Per part number matters: mixing part families averages away the step you are looking for.
- Pull goods-receipt records for the input materials over the same period, with the timestamp at which each lot was first issued to the line — not the date it arrived in the yard.
- Plot them on one time axis. You are looking for a step, not a slope, and for whether the step edge lands within an hour or two of a changeover.
- Check the counter-example. If the previous three changeovers produced no step, the material is not automatically the variable — it may be this lot in particular, which is a supplier conversation rather than a model one.
- Note what else happened that morning. Shift start, a lamp replacement, a fixture change and a lot changeover often coincide, and only one of them is the cause.
If the step does not align with a lot boundary, look elsewhere before you go further. A reject rate that follows the clock rather than the goods-receipt log is an illumination problem, diagnosed in the model that finds the scratch on day shift and misses it after dark.
Worse, or merely different: the 30-part retest
Take 30 consecutive rejects from the new lot to the inspection bench and have them judged against the written class list by someone who does not know why they are being asked. Count how many the human confirms as genuinely non-conforming. That single number decides everything that follows, and it takes under an hour.
- Nearly all confirmed. The lot is worse. The system found it, the escape risk was real, and the conversation is with purchasing and the supplier — not with engineering.
- Nearly none confirmed. The lot looks different and remains good. This is overkill, it is expensive, and the fix is revalidation of the inspection against the new appearance.
- A genuine mix. Usually two things at once: a real quality shift plus an appearance shift. Split the rejects by defect class before deciding, because one class may be real and the rest cosmetic.
- Disagreement between two inspectors on the same parts. Stop and fix that first — a class list two people read differently cannot arbitrate anything, as labelling defect images so two inspectors agree sets out.
The four properties that move a score distribution
| Property | What the camera sees | How to confirm | Response |
|---|---|---|---|
| Specular gloss | Brighter highlights and saturated pixels on a bright-field station; a lifted background on a dark-field one | Gloss readings on both lots with the same instrument and the same geometry | Re-tune illumination angle or polarisation before touching the model |
| Colour or pigment load | A shift of tens of grey levels in a monochrome image from a difference the eye barely registers | Photograph both lots under the station's own lighting and compare grey-level histograms | Consider a filter or a colour channel, or accept a per-lot baseline |
| Coating or plating thickness | Changed reflectance and, on thin films, visible banding that the model reads as a surface flaw | Thickness measurements from the incoming certificate, cross-checked on a sample | Add the appearance range to the material specification, not to the training set |
| Rolled, brushed or moulded finish | A different background texture, which for a normality model is a different definition of normal | Compare texture statistics on defect-free areas of both lots | Revalidate; a new supplier's finish is a new domain, not more data |
The second row is the one that surprises people most. A monochrome sensor's spectral response is not the eye's, and neither is the station's illumination, so two lots a colour inspector signs off as identical can land 20 or 30 grey levels apart, out of 255, in the image the model actually receives. If your defect judgement rests on absolute brightness anywhere in the pipeline, colour drift will find it. An 8-bit pipeline has no headroom to hide that in.
A model trained on one lot has learned what that lot looks like, and a new lot is a new question. Widening the definition of normal until it accommodates both is not robustness, it is a slow surrender of the detection you paid for.
Lot-aware thresholds, and when they are the wrong answer
The tempting fix is a per-lot threshold: recalibrate on the first parts of each lot and carry on. It works, in a bounded set of circumstances, and it is dangerous outside them. Recalibrating against the incoming population assumes that population is mostly good; if the lot carries a systematic defect, an auto-tuned threshold quietly learns to accept it and the escape rate goes up while the dashboard looks better than ever.
The alternative to per-lot thresholds is per-lot revalidation with a fixed threshold — slower, more disruptive, and much easier to defend when a customer asks how the accept criterion was set on the day their part was made.
The standing control: a revalidation run on first use of a lot
- Trigger it from goods receipt, not from a person noticing. The first issue of a new lot to a line should raise the run automatically — the kind of small event-driven automation work that is worth building once and forgetting.
- Run 50 parts under normal conditions and capture every score, not just the dispositions.
- Compare the score distribution against the running baseline for that part number: the median, the spread, and the 95th percentile against the threshold.
- Hand-inspect every reject in the run, and a random sample of 10 passes. Record agreement between the human and the system as a number.
- Check the illumination proxy at the same time — the background grey-level statistic — so a lighting change is not misread as a material change.
- Exit with one of three states: release, release under an operator gate with elevated sampling, or quarantine the lot for engineering. Write the state, the evidence and the owner into the lot record.
Fifty parts and an hour is the price. It is small compared with a shift of overkill, and much smaller than the alternative of finding out from a customer. Where a lot fails the run, resist the reflex to retrain immediately: capture the images, keep the lot separate, and decide whether the correct answer is a lighting change, a fixture change, a specification conversation with the supplier, or genuinely new training data. Presentation is often the cheapest lever, which is why fixturing a part so the camera sees the same thing twice is worth reading before the training set is opened.
The traceability that makes any of this possible
Every method on this page assumes one thing: the lot identifier is attached to the inspection record at capture. If it is not — and in most plants it is not — you cannot overlay anything, because the images know the time and the camera but not the material. The lot number lives on a printed traveller, gets copied by hand into a spreadsheet at the end of the shift, and never reaches the system that took the picture. That gap is the subject of why the floor keeps printing the paper copy, and closing it is a prerequisite rather than an improvement.
- Stamp lot, supplier, goods-receipt reference and issue time onto every image record, alongside model version and threshold version.
- Keep the score even when you discard the image. Score histories are tiny and they are what the overlay is built from.
- Retain enough per-lot history to establish a baseline. 1 previous lot is an anecdote; 6 give you a spread you can quote a tolerance against.
- Keep the corpus where the plant can use it. Revalidation and retraining both need images, scores and the model in the same place, which is why plants that will not send video off site reach the same conclusion as teams weighing running models on their own infrastructure.
- If the answer to a lot change is a second camera or a higher resolution, check the timing budget before committing — the constraints in it keeps up on the bench and falls behind at line speed apply to every added view.
- And rule out acquisition first if images look wrong rather than merely different, using half a part in frame: trigger and encoder faults.
Treated properly, incoming material becomes a first-class variable in the inspection system rather than an unmodelled input that occasionally ruins a Tuesday. The rest of the decisions around it sit in visual inspection and defect detection, part of our manufacturing and industrial vision work.
Frequently asked questions
Short answers to the follow-ups this page tends to raise.
Why did our vision system start rejecting good parts after a new material batch?
Because the new batch looks different from the one the system was set up on, and appearance is what it measures. Gloss, pigment, coating thickness and surface finish all shift the score distribution for every part at once, so a threshold that used to sit clear of the pass population now cuts through it. Confirm by overlaying reject rate on goods-receipt changeover times and looking for a step at the boundary.
How do we know whether the material is worse or just different?
Have 30 consecutive rejects judged by hand against the written defect class list. If a human confirms most of them, the lot genuinely is worse and the system is working — take it to the supplier. If a human passes most of them, the lot is good and merely looks different, which is a revalidation job on the inspection rather than a quality problem in the material.
Should we retrain the model on the new lot?
Not as the first move, and never on unverified parts. Retraining on a lot you have not inspected teaches the model to accept whatever that lot contains, including defects, and each such round widens the definition of normal a little further. Revalidate first, fix presentation or illumination if that is the cause, and add training data only when the new appearance is genuinely a legitimate new normal.
Is a per-lot threshold a reasonable solution?
Only under conditions you should state explicitly. The defect has to be rare enough that incoming parts are mostly good, the baseline has to come from a run a human has checked, the adjustment has to be bounded so a large shift stops the line instead of recalibrating, and every change has to be logged against the lot. Without those four, an auto-tuning threshold will eventually calibrate itself onto a bad lot.
- material variation
- inspection
- revalidation
- quality
The work behind this page
Builds from our portfolio that this page draws on.
Open Vision PPE Monitoring
Boundary surveillance, PPE compliance monitoring, and intrusion detection via real-time video analytics. Runs fully on-premise — no cloud required.
Safety & ComplianceFactory OS
Production planning and task management for a tier-1 apparel manufacturer — replacing Excel with automated milestone planning, SOP gate enforcement, and real-time visibility.
ManufacturingRead next
- The model finds the scratch on day shift and misses it after darkA detection rate that tracks the clock is an optics fault wearing a model's clothes. Re-scoring yesterday's archived images tonight proves the model is unchanged and points the investigation at the light instead.diagnostic
- Half a part in frame: trigger and encoder faults that look like the modelStreaked surfaces, clipped parts and duplicated frames are acquisition faults with physical causes and arithmetic answers. Counting triggers against parts for 100 pieces tells you which one you have.diagnostic
- The camera passed it and the customer found it: tracing an escape backwardsMost escape investigations start by retraining a model that never saw the surface in question. Working backwards through the retrieval chain — part to image to score to disposition — settles which of five very different faults you actually have.diagnostic
- False reject rate: what the line feels, not the accuracy on the slideA 1% false reject rate sounds like rounding. At 1,200 parts an hour it is 96 good parts a shift and over an hour of somebody re-checking them.definition
- It keeps up on the bench and falls behind at line speedA cell that is fast on the bench and late on the line has either a latency problem or a throughput problem, and they need opposite fixes. Queue depth over ten minutes tells you which, and a per-part budget tells you where the time went.diagnostic
- The anomaly score is a ranking; the threshold is a business decisionThe score orders parts by how far they sit from normal. It carries no units, no probability and no severity — which is why the threshold belongs to quality, not to engineering.definition
Working on something in this space?
Tell us where you are in a sentence or two. We'll tell you honestly whether we're the right team, and what a sensible first slice of the work looks like.
Start the conversation