ML//neural network//autoencoder

An autoencoder is a neural network trained to copy its input to its output through a narrow bottleneck, and its main industrial use is anomaly detection: trained only on healthy data, it learns what normal looks like and copies anything else badly. An encoder \(g\) compresses the input \(x\) into a small vector \(z=g(x)\) and a decoder \(f\) tries to rebuild \(x\) from it. Training minimizes the reconstruction error on normal data, and in service that error becomes the anomaly score.


An autoencoder is a neural network trained to copy its input to its output through a narrow bottleneck, and its main industrial use is anomaly detection: trained only on healthy data, it learns what normal looks like and copies anything else badly. An encoder ggg compresses the input xxx into a small vector z=g(x)z=g(x)z=g(x) and a decoder fff tries to rebuild xxx from it. Training minimizes the reconstruction error on normal data, and in service that error becomes the anomaly score.

s(x)=∥x−f(g(x))∥2s(x)=\left\lVert x-f\big(g(x)\big)\right\rVert^2s(x)=​x−f(g(x))​2

A vibration window from a healthy pump reconstructs with a small sss; a window carrying a bearing defect does not fit the patterns the bottleneck learned and comes back distorted. The threshold sits at a high percentile of the scores on validation data (the 99.5th, for instance), better set per operating regime, since a pump at full flow and at idle have different normals.

A counterfeiter who has only ever seen genuine banknotes.

Hand it something odd and its copy comes out wrong; the difference between the original and the copy is the alarm.

A linear autoencoder is PCA.

With linear encoder and decoder it learns the same subspace as PCA, and its score is exactly the Q statistic. The autoencoder is a nonlinear PCA, so PCA goes first: if it suffices, the network is saved.

It pays when normal data lies on a curved surface that a plane cannot follow, as with vibration windows or images, and when there are many more healthy samples than PCA would need; the price is interpretability, since a high score no longer says which variables broke the pattern. The order of escalation is a threshold per variable, then multivariate statistical process control, then the network (flyswatter rule).

It fails in three known ways. With too much capacity it learns to copy everything, anomalies included, so the bottleneck must stay narrow; unlabeled failures in the training data are learned as normal; and any legitimate change of operation (a new product, a new setpoint) raises the alarm until it is retrained.

The same architecture has other uses with other names: the sparse autoencoder reads features out of a language model's activations, and the VAE makes the bottleneck probabilistic so it can generate new samples. As a detector it is one member of anomaly detection, and its score is a residual like any other (residual as surprise).