mathematics//statistics//effective sample size
The effective sample size is the number of independent observations that would carry the same information as a set of correlated or unequally weighted ones, and it is the honest count to use whenever samples have memory or weights: a fast-logged sensor, a particle filter's cloud, the output of an MCMC chain. Error bars, tests and data requirements computed with the raw row count are wrong by the ratio between the two.
The effective sample size is the number of independent observations that would carry the same information as a set of correlated or unequally weighted ones, and it is the honest count to use whenever samples have memory or weights: a fast-logged sensor, a particle filter's cloud, the output of an MCMC chain. Error bars, tests and data requirements computed with the raw row count are wrong by the ratio between the two.
A slowly varying signal shows it best. When each sample repeats most of the previous one, with correlation ρ\rhoρ between consecutive samples (a first-order process, autocorrelation),
Neff≈N 1−ρ1+ρ.N_{\text{eff}}\approx N\,\frac{1-\rho}{1+\rho}.Neff≈N1+ρ1−ρ.
An hour logged at 1 kHz is 3.6 million samples; if the signal has a second of memory, ρ≈0.999\rho\approx0.999ρ≈0.999 and the hour is worth about 1,800 independent readings. A historian full of such data is much less information than its size suggests (historian).
Count effective samples, never rows.
Every formula that divides by N\sqrt NN<path d="M95,702
c-2.7,0,-7.17,-2.7,-13.5,-8c-5.8,-5.3,-9.5,-10,-9.5,-14
c0,-2,0.3,-3.3,1,-4c1.3,-2.7,23.83,-20.7,67.5,-54
c44.2,-33.3,65.8,-50.3,66.5,-51c1.3,-1.3,3,-2,5,-2c4.7,0,8.7,3.3,12,10
s173,378,173,378c0.7,0,35.3,-71,104,-213c68.7,-142,137.5,-285,206.5,-429
c69,-144,104.5,-217.7,106.5,-221
l0 -0
c5.3,-9.3,12,-14,20,-14
H400000v40H845.2724
s-225.272,467,-225.272,467s-235,486,-235,486c-2.7,4.7,-9,7,-19,7
c-6,0,-10,-1,-12,-3s-194,-422,-194,-422s-65,47,-65,47z
M834 80h400000v40h-400000z"/> (a standard error, a confidence interval, a test) needs the effective count; with the raw count, the error bar on that hour of data comes out about 45 times too narrow.
Weights shrink it too. A particle filter with weights w(i)w^{(i)}w(i) summing to one has Neff=1/∑i(w(i))2N_{\text{eff}}=1/\sum_i\left(w^{(i)}\right)^2Neff=1/∑i(w(i))2: equal weights give NNN, one particle holding all the weight gives 1. The usual rule is to resample when it drops below N/2N/2N/2, before the cloud degenerates into a few survivors.
MCMC samples are autocorrelated by construction, so a chain of 10,000 draws may be worth a few hundred; diagnostic tools report the effective size of each parameter, and that number decides whether the run was long enough (MCMC).
Matrix statistics feel it as well. The noise band of eigenvalues that Marchenko-Pastur law predicts for a sample covariance assumes independent rows; with autocorrelation the true NNN is smaller and the band wider, so more components look like signal than are.
Decimating does not create information. Keeping every hundredth sample of the 1 kHz log loses almost nothing, which is the practical way to make standard tools that assume independence behave; filtering and then downsampling is cheaper than pretending the rows are independent.