ML//unsupervised learning//clustering
Clustering is the family of unsupervised learning methods that split unlabelled samples into groups whose members resemble each other more than they resemble the rest, and in industry it is used mainly to discover structure nobody wrote down: the operating regimes of a compressor, families of pumps with similar vibration signatures, customer segments. Two hundred pumps, each described by its vibration spectrum and its consumption, fall into four groups without anyone labelling a single one, and each group then gets its own normal behaviour and its own alarm thresholds (operating regime).
Clustering is the family of unsupervised learning methods that split unlabelled samples into groups whose members resemble each other more than they resemble the rest, and in industry it is used mainly to discover structure nobody wrote down: the operating regimes of a compressor, families of pumps with similar vibration signatures, customer segments. Two hundred pumps, each described by its vibration spectrum and its consumption, fall into four groups without anyone labelling a single one, and each group then gets its own normal behaviour and its own alarm thresholds (operating regime).
Every method in the family needs two decisions before any algorithm runs: what makes two samples similar (a distance, a kernel, a graph of neighbours, all sensitive to how the features are scaled) and what counts as a group (compact around a centre, connected through neighbours, dense against a sparse background). The members differ in that second choice, which is why the same data can yield different groups under each.
A clustering always returns the groups it was asked for, whether or not they exist.
Four clusters do not prove four operating modes; that is checked against the physics or an operator who knows the plant. And if the mode is already in the PLC record (the recipe, the speed setpoint), grouping by that column beats any algorithm.
K-means places KKK centres and assigns each sample to the nearest, alternating with moving each centre to the mean of its samples. It is fast, runs in production everywhere and finds roundish groups of similar size; it needs KKK given in advance.
A Gaussian mixture model softens K-means: each group is a Gaussian with its own mean and covariance matrix, fitted by expectation-maximization, so groups can be elongated ellipses and each sample gets a probability of belonging to each. It is the better choice when regimes overlap or are tilted in feature space, at a little more cost and with the same need to fix KKK.
Spectral clustering builds a graph of similarities between samples and runs K-means on the slow eigenvectors of its graph Laplacian, so rings, bands and regimes that curve through measurement space become compact groups. It costs an eigendecomposition of an N×NN\times NN×N matrix, manageable with thousands of samples, and its result depends on the similarity scale.
Groups found this way describe the data and prove no cause; using them as states of a model or as labels for a classifier needs its own check (sufficient state).