Defect Spotter

Unsupervised visual inspection · PatchCore (CVPR 2022)

It has never seen a defect.
It still finds them.

Defect Spotter learns only from photos of good parts. Anything that doesn't look like those photos (a crack, a missing component, a stain) lights up on a pixel-level heatmap. No labelled defects needed, which is exactly what factories don't have.

-mean image AUROC
12 VisA categories
-mean pixel AUROC
defect localisation
-per image
on a 6-core CPU, no GPU
0defect images
used for training

Benchmark explorer

Results on the official VisA test split. Heatmaps are the model's output; green outlines are the human-labelled ground truth.

-

What "normal" looks like (training images):

Where to draw the line?

Every inspection system trades missed defects against false alarms. Drag the threshold: the numbers are recomputed on the real test scores. The marker shows the deployed threshold (5% false-alarm budget, set on good parts only).

Test your own image against this model

Upload any photo. The model only knows this category, so other objects will (correctly) look anomalous.

Teach it your own object, live

Take 5–20 photos of something intact (a mug, a coin, a keyboard, a sticker) with your phone or webcam. The model is built in seconds. Then photograph it with a scratch, a stain or a missing piece.

1

Show it normal

Same object, similar framing and background. A little variation in angle helps.

0 photos

2

Build the model

Builds a memory bank of patch features, then sets the threshold by leave-one-out on your own photos.

3

Inspect

Now photograph it with a defect, or a normal one to check it stays quiet.

Privacy: photos are processed in memory and discarded. Nothing is saved, and sessions expire after 30 minutes.

How it works

No training loop and no labels: it's nearest-neighbour search in a well-chosen feature space.

Image256→224 px Frozen CNNWideResNet-50 · layers 2+3 Patch features28×28 × 1024-d, 3×3 context Memory banknormal patches, 5% coreset Nearest neighbourdistance per patch Heatmapmax = image score

Train

Pass ~200 good images through an ImageNet CNN, keep every local patch descriptor, then pick the 5% that best cover the space (greedy k-center coreset).

Calibrate

The alarm threshold is set on held-out good images (99th percentile + margin), never on defects, so the reported accuracy is honest.

Detect

Each test patch's distance to its nearest normal patch is its anomaly score. Upsampled and smoothed, these scores become the heatmap.

Limits

It flags unusual, not defective: new lighting or an unseen but acceptable variant can trigger it. Pose and background must match the training photos.