industrial//process monitoring//multivariate statistical process control
Multivariate statistical process control is a monitoring method that learns how a plant's variables normally move together and flags every sample that breaks that pattern, using PCA fitted on healthy operation; it is the oldest form of anomaly detection in process industry (chemicals, pharmaceuticals, semiconductors) and still the one that runs in most control rooms. It needs weeks of clean normal data covering the operating modes to be watched, and no labelled failures.
Multivariate statistical process control is a monitoring method that learns how a plant's variables normally move together and flags every sample that breaks that pattern, using PCA fitted on healthy operation; it is the oldest form of anomaly detection in process industry (chemicals, pharmaceuticals, semiconductors) and still the one that runs in most control rooms. It needs weeks of clean normal data covering the operating modes to be watched, and no labelled failures.
The recipe has two phases. Offline, the variables are standardized and PCA is fitted on the normal weeks, keeping kkk components: the plane of normal operation, spanned by VkV_kVk, with variances λi\lambda_iλi. Online, each new standardized sample xxx gets its coordinates in that plane, t=VkTxt=V_k^{\mathsf T}xt=VkTx, and two statistics:
T2=∑i=1kti2λi,Q=∥x−VkVkTx∥2.T^2=\sum_{i=1}^{k}\frac{t_i^2}{\lambda_i},\qquad Q=\bigl\|x-V_kV_k^{\mathsf T}x\bigr\|^2 .T2=i=1∑kλiti2,Q=x−VkVkTx2.
The T-squared statistic is the distance inside the plane, measured in standard deviations: the plant is at an unusual point but every relation holds. The Q statistic is the squared distance to the plane: the correlations that physics imposed have broken.
PC1 / PC2 variance95 % / 5 % anomalies caught by Q8 of 8 correlation r0.88 PCA fitted on 250 normal readings of two correlated sensors: PC1 explains 95 % of the variance. Eight injected readings sit near the centre but off the PC1 axis, so neither sensor leaves its range; the residual Q, their squared distance to the axis, catches 8 of 8 above the 99 % threshold of the normal readings. Multiplying sensor B by 10 turns PC1 to 83° from A and makes it explain 99.6 %, through units alone; standardizing brings it back to 94 %.
The eight injected readings sit near the centre, inside both sensors' normal ranges, and only Q flags them; drag a reading along the axis and then off it, and only the second move sets Q off, and scale sensor B by 10 to see the plane swing toward it until standardizing puts it back.
PCA finds the natural axes of the plant's data, and the two statistics split what leaves them.
T-squared watches what is unusual but coherent; Q watches what breaks the physics. A fault that only a combination of normal values reveals is caught at the cost of one matrix-vector product per sample, small enough for a PLC.
Control limits for T-squared and Q come from classical approximations (an F distribution for T-squared, the Jackson and Mudholkar formula for Q) or, more robustly, from the 99th percentile of each statistic on normal data that were not used for the fit.
A contribution plot, which variables add most to an alarming Q, gives the first hint of where to look. It is the start of fault isolation, not its end: a fault spreads through the balances and lights up its neighbours too.
The method is linear, so strongly curved relations fill Q with false alarms. It assumes one regime, so several operating regimes need one model each. It ignores dynamics, so autocorrelated samples call for dynamic PCA, which adds lagged copies of the variables as extra columns. And it ages: after major maintenance the normal plane moves and the model is refitted.
It belongs to the wider practice of process monitoring, where it is the step after per-variable limits.