mathematics//random matrix theory//Marchenko-Pastur law

The Marchenko-Pastur law is the result of random matrix theory that gives the distribution of the eigenvalues of a sample covariance computed from pure noise, and it is used as the yardstick against which principal components, factors and modes are judged: only what clearly rises above the noise band is a candidate for signal. If \(p\) variables are independent noise of variance \(\sigma^2\) and the covariance is estimated from \(N\) samples (with \(p\le N\)), the eigenvalues do not all come out at \(\sigma^2\); for large matrices they fill the interval


The Marchenko-Pastur law is the result of random matrix theory that gives the distribution of the eigenvalues of a sample covariance computed from pure noise, and it is used as the yardstick against which principal components, factors and modes are judged: only what clearly rises above the noise band is a candidate for signal. If ppp variables are independent noise of variance σ2\sigma^2σ2 and the covariance is estimated from NNN samples (with p≤Np\le Np≤N), the eigenvalues do not all come out at σ2\sigma^2σ2; for large matrices they fill the interval

λ±=σ2(1±p/N)2.\lambda_{\pm}=\sigma^2\left(1\pm\sqrt{p/N}\right)^2 .λ±​=σ2(1±p/N​)2.

λ−\lambda_-λ−​ and λ+\lambda_+λ+​ are the edges of the band where noise eigenvalues fall, and only the ratio p/Np/Np/N sets its width. With 500 sensors and 1,000 samples, 0.5≈0.71\sqrt{0.5}\approx0.710.5​≈0.71 and the band runs from 0.09σ20.09\sigma^20.09σ2 to 2.91σ22.91\sigma^22.91σ2. A component of pure noise with almost three times the average variance is therefore what noise alone produces. The check takes three lines of NumPy: draw 1,000 rows of 500 Gaussian columns, take np.linalg.eigvalsh of their covariance, and the largest eigenvalue comes out near 2.9.

Before interpreting a component, compare it with Marchenko-Pastur.

Estimate σ2\sigma^2σ2 from the bulk of the eigenvalues, compute λ+\lambda_+λ+​, and treat as candidate signal only the components that stand clearly above it; the rest is the shape of the noise.

It assumes independent samples. Autocorrelated data, as any sensor sampled faster than its process moves, carry fewer independent samples than rows, so the effective sample size is smaller than NNN and the true band is wider than the formula with the raw NNN says.

It assumes equal variances across variables. Mixing pressures in pascals with temperatures in degrees breaks that at once, so the data are standardized first, as PCA on plant data requires anyway.

It sets a floor for detection as well as a ceiling for noise. A real but weak signal can sit below λ+\lambda_+λ+​ and be invisible to any eigenvalue test at that sample size; the only remedy is more independent data.

The same yardstick guards the monitoring models built on PCA: a multivariate statistical process control model that keeps components inside the noise band spends its T2T^2T2 statistic on noise, and choosing how many to keep by comparison with λ+\lambda_+λ+​ is more honest than a fixed share of explained variance.