How Does a Machine “See” a Defect? Computer Vision Explained Without the Magic
Visual inspection is where AI meets pharma most concretely: cameras, vials, blisters, labels moving at hundreds of pieces per minute. But when I ask people how they think the machine recognizes a crack, the answers reveal a fundamental misunderstanding worth clearing up.
A digital image, for a computer, is just a grid of numbers – pixel intensities. Traditional machine vision works by having engineers define explicit operations on those numbers: thresholds, edge filters, geometric measurements. “If the dark region in this zone exceeds N pixels, reject.” This works beautifully for stable, well-defined tasks – dimensional checks, presence/absence, code reading – and it remains the right tool for many of them.
Deep-learning-based vision works differently. A convolutional neural network is shown thousands of labeled images – good products and defective ones – and it learns, layer by layer, its own internal representations: early layers detect edges and textures, deeper layers combine them into shapes, and the final layers learn “this combination of patterns means a crack near the shoulder of the vial.” Nobody programs those representations; they emerge from training.
The practical consequences are what matter:
- Strength: the network handles natural variability – lighting shifts, product orientation, cosmetic differences – far better than hand-tuned rules, because it has seen variability during training. This is why AI shines on difficult defects: scratches vs. fibers, cosmetic vs. functional anomalies.
- Weakness: the network only knows what it has seen. A brand-new defect type, a new container format, a changed light source – these can degrade performance silently. This is why dataset management, periodic revalidation, and retraining workflows are not optional extras: they are the system.
- Hybrid reality: in serious industrial deployments, deep learning rarely replaces classical vision entirely. It complements it – rules for the deterministic checks, neural networks for the perceptual judgment calls, and often a human review loop for borderline cases that feed the next retraining cycle.
Understanding this demystifies the “magic” and clarifies where responsibility lies: not in the algorithm, but in the quality of the examples you teach it with.