Manufacturing and Predictive Maintenance
Visual defect inspection
Visual defect inspection compares every unit against what a good one is supposed to look like, catching scratches and flaws by how much a product deviates from that reference.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Visual defect inspection compares every product against what a good one is supposed to look like, and flags anything that deviates too much.
Think about a baker glancing at every biscuit coming off the conveyor belt, pulling out the ones that are burnt, cracked, or oddly shaped. The baker is not consciously listing rules — "reject if the edge colour is darker than X." They have seen thousands of good biscuits. Anything that looks meaningfully different jumps out immediately.
A camera-based inspection system on a factory line does the same job. It works at far higher speed and consistency than a tired human eye at the end of a long shift.
Why it exists
Manufacturing defects — a scratch, a dent, a misprint, a crack — are usually rare, and each one can look different from the last. Training a model the ordinary way means showing it thousands of labelled examples of every possible defect. That runs into the same wall covered a couple of lessons ago. There may only be a handful of confirmed defective units on file. That is nowhere near enough to teach a model every way a product can go wrong.
The practical fix flips the problem around. Instead of trying to teach a model every possible defect, teach it very well what a good unit looks like. That part is easy, because good units are the overwhelming majority of what a factory produces. Anything that deviates significantly from that learned "normal" gets flagged, whether or not that specific type of flaw has ever been seen before.
How it works
Step 1: Photograph many known-good units under the same lighting and angle.
Step 2: Build a "golden template" -- what a normal surface should look like,
pixel by pixel, including how much natural variation is expected.
Step 3: For each new unit:
compare its image against the golden template, pixel by pixel.
A good unit: small, expected differences everywhere
A flawed unit: one region differs FAR more than expected
-> flagged and pulled from the lineThis is sometimes called golden template matching, a real, widely used technique in automated optical inspection. It needs no labelled defect examples at all to get started.
Where you have already seen it
- Printed circuit board inspection, where cameras check every board against a reference layout before it leaves the factory.
- Fruit sorting machines, which photograph each fruit and reject ones with bruises or discoloration that a healthy fruit's surface would not have.
- Automated bottling line cameras, checking fill levels and cap seating against what a correctly filled, sealed bottle should look like.
An honest warning
A vision-based inspection system is only as good as its lighting, camera angle, and the range of "normal" variation it was shown during setup. A perfectly good unit photographed under unusual lighting can be wrongly flagged. A genuinely new type of defect, one the system was never calibrated to notice, can slip through unnoticed. These systems assist a quality process — they do not remove the need for a defined escalation path when something looks uncertain.
Remember this
- Defects are rare and varied, so it is often easier to teach a model what "normal" looks like than to teach it every possible flaw.
- Comparing a new unit against a learned reference of "good," pixel by pixel, catches unfamiliar defects too, not only ones seen before.
- Camera setup, lighting, and the range of natural variation all directly affect how reliable this kind of system actually is.
What to learn next
- Soft sensors — estimating a quantity a camera or sensor cannot measure directly.
- Building a model from four failures — the same "learn normal, flag deviation" idea, applied to sensor time series instead of images.
- What is computer vision? — the general foundation this lesson's technique builds on.
Developer — Code and libraries.
Setup
pip install numpyMinimal runnable code
We simulate a product's surface as a small textured grayscale image, build a "golden template" from 30 known-good units, and check whether a simple per-pixel comparison can catch a scratch on a new unit.
import numpy as np
rng = np.random.default_rng(23)
size = 20 # a small 20x20 "image" of a product's surface, in grayscale
def make_good_surface():
# A consistent, repeatable base pattern (like a moulded or printed
# surface) plus small camera and lighting noise.
x = np.linspace(0, 3 * np.pi, size)
base = 128 + 20 * np.sin(x)[:, None] * np.cos(x)[None, :]
return base + rng.normal(0, 3, (size, size))
def add_scratch(image):
flawed = image.copy()
flawed[8:12, 5:15] += 60 # a bright scratch-like streak
return flawed
# Build a "golden template" from many known-good units -- exactly what a
# real automated optical inspection (AOI) system does at setup time.
n_reference = 30
reference_images = np.stack([make_good_surface() for _ in range(n_reference)])
golden_mean = reference_images.mean(axis=0)
golden_std = reference_images.std(axis=0) + 1e-6 # avoid dividing by zero
print(f"golden template built from {n_reference} known-good units")
print(f"typical per-pixel noise (std) in a healthy unit: {golden_std.mean():.2f}")
print()
def anomaly_score(image, golden_mean, golden_std):
z_scores = np.abs(image - golden_mean) / golden_std
return z_scores.max(), (z_scores > 5).sum() # worst pixel, and count of very abnormal pixels
good_unit = make_good_surface()
flawed_unit = add_scratch(make_good_surface())
good_score = anomaly_score(good_unit, golden_mean, golden_std)
flawed_score = anomaly_score(flawed_unit, golden_mean, golden_std)
print(f"good unit -> worst pixel z-score: {good_score[0]:.1f}, pixels flagged: {good_score[1]}")
print(f"flawed unit -> worst pixel z-score: {flawed_score[0]:.1f}, pixels flagged: {flawed_score[1]}")
print()
# Test on 200 more units (10 with a scratch) and see how well a simple
# threshold on "pixels flagged" separates them.
n_test = 200
n_flawed = 10
flags = []
truth = []
for i in range(n_test):
is_flawed = i < n_flawed
unit = add_scratch(make_good_surface()) if is_flawed else make_good_surface()
_, n_flagged = anomaly_score(unit, golden_mean, golden_std)
flags.append(n_flagged)
truth.append(is_flawed)
flags = np.array(flags)
truth = np.array(truth)
threshold = 5
predicted_flawed = flags > threshold
print(f"threshold: more than {threshold} abnormal pixels = reject the unit")
print(f"true flawed units caught: {(predicted_flawed & truth).sum()} / {truth.sum()}")
print(f"good units wrongly rejected: {(predicted_flawed & ~truth).sum()} / {(~truth).sum()}")golden template built from 30 known-good units typical per-pixel noise (std) in a healthy unit: 2.93 good unit -> worst pixel z-score: 4.2, pixels flagged: 0 flawed unit -> worst pixel z-score: 30.1, pixels flagged: 40 threshold: more than 5 abnormal pixels = reject the unit true flawed units caught: 10 / 10 good units wrongly rejected: 0 / 190
What actually happened
golden_mean and golden_std describe, pixel by pixel, what a healthy unit's surface normally looks like and how much natural variation to expect there — built entirely from good units, no defect examples needed.
anomaly_score computes a z-score for every pixel: how many standard deviations away from the expected value that pixel is. A healthy unit's worst pixel scored 4.2 — a bit unusual, but within the range ordinary camera and lighting noise can produce. The scratched unit's worst pixel scored 30.1, wildly beyond anything normal variation explains, and 40 separate pixels crossed the abnormal threshold.
On the larger test of 200 units, a simple rule — "more than 5 abnormal pixels means reject" — caught every flawed unit and wrongly rejected none of the good ones. Real production inspection is rarely this clean; this result reflects a synthetic example with a large, obvious defect, not a guarantee for subtle real-world flaws.
Common mistakes
Building the golden template from too few reference units. With too little data, golden_std underestimates real natural variation, and the system rejects perfectly good units constantly — a very expensive kind of false alarm on a production line.
Assuming lighting and camera position stay constant forever. A camera that shifts slightly, or a light bulb that dims with age, changes what "normal" looks like — and needs the golden template rebuilt, or the system starts rejecting good units for the wrong reason.
Using one global threshold when a product has naturally uneven surfaces. A product with an intentional pattern, texture or printed logo needs golden_std to vary by region, exactly as this example already does per-pixel — a single flat threshold across the whole image would misfire constantly on the busier, naturally varied parts of the surface.
Treating this as a full defect classifier. This method answers "does this look abnormal," not "what specific defect is this." A real quality system typically pairs this anomaly step with a further classification step, once enough labelled defect examples exist to train one reliably.
Try it yourself
Change the scratch's added brightness from += 60 to += 10, simulating a much fainter, subtler defect. Re-run the 200-unit test and see how many flawed units the same threshold now misses — a concrete demonstration of why defect severity, not only presence, changes how hard detection really is.
What to learn next
Researcher — Mathematics and papers.
From golden templates to learned reconstruction
The pixel-wise z-score approach above is a classical, interpretable baseline. It assumes good units are well aligned (same position, angle, scale) and that natural variation is well described by an independent per-pixel Gaussian — both assumptions that break down for products with natural shape variation, or images that are not perfectly registered against the template.
Autoencoder-based anomaly detection generalises the same "learn normal, flag deviation" idea without those assumptions. An autoencoder is trained to compress and then reconstruct only images of good units. At inference time, a defective region reconstructs poorly — the model has never learned to represent it — producing a large, spatially localised reconstruction error exactly where the defect sits:
anomaly_map(x) = | x - decoder(encoder(x)) |x— the input imageencoder,decoder— the trained autoencoder's two halvesanomaly_map— a per-pixel error map, high where the model failed to reconstruct the input, which tends to coincide with defective regions
Bergmann et al. (2019), MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection, CVPR, established the standard public benchmark for this task, with real industrial images across many product categories and defect types, and demonstrated that reconstruction-based methods substantially outperform simple statistical baselines once product geometry and texture become complex.
Feature-embedding approaches
More recent state-of-the-art methods (PaDiM, Defard et al., 2021; PatchCore, Roth et al., 2022) skip pixel-level reconstruction entirely. They extract deep feature embeddings from a pretrained convolutional network for patches of known-good images, model the distribution of those embeddings (a multivariate Gaussian per spatial location, or a memory bank of representative "normal" patch embeddings), and score a new image by how far its patch embeddings sit from that learned normal distribution. This sidesteps the alignment and reconstruction-fidelity issues that limit both the golden-template and autoencoder approaches, at the cost of needing a pretrained backbone and more computation per image.
Why unsupervised framing dominates this application specifically
Supervised object detection or segmentation, trained directly on labelled defect images, generally outperforms unsupervised anomaly detection when enough labelled defects exist to train it properly — often several hundred to a few thousand per defect type. In practice, that labelled volume rarely exists early in a product's life, and defect types keep appearing that were never seen during training. This is exactly the same rare-event data problem covered in Building a model from four failures, applied to images instead of sensor time series, and it is why unsupervised or few-shot approaches dominate the vision-inspection literature specifically, even though supervised methods are the default choice for well-labelled vision tasks elsewhere.
Evaluation in production
A defect-inspection system is evaluated with the same cost-asymmetric thinking as Choosing a threshold from costs: missing a real defect (a false negative, or "escape") typically costs far more than rejecting a good unit (a false positive, or "false reject"), since an escaped defect can reach a customer. Production systems are usually tuned toward high recall on defects, with the resulting false-reject rate managed through a secondary human review station rather than accepted as pure waste.
Key references
- Bergmann, P., Fauser, M., Sattlegger, D. & Steger, C. (2019). MVTec AD: A Comprehensive Real-World Dataset for Unsupervised Anomaly Detection. CVPR.
- Defard, T., Setkov, A., Loesch, A. & Audigier, R. (2021). PaDiM: A Patch Distribution Modeling Framework for Anomaly Detection and Localization. ICPR.
- Roth, K. et al. (2022). Towards Total Recall in Industrial Anomaly Detection. CVPR (PatchCore).