mathematics//statistics//goodness of fit

Whether a sample could have come from a given distribution, most often the normal distribution, is a question of **goodness of fit**. It depends on the data collected, and it is answered first with a picture and only then with a test.


Whether a sample could have come from a given distribution, most often the normal distribution, is a question of goodness of fit. It depends on the data collected, and it is answered first with a picture and only then with a test.

The picture is the QQ plot (quantile against quantile). Sort the standardized data, and for the iii-th value compute where a normal would put the same rank, its theoretical quantile Φ−1 ⁣((i−3/8)/(n+1/4))\Phi^{-1}!\big((i-3/8)/(n+1/4)\big)Φ−1((i−3/8)/(n+1/4)). Plot one against the other: if the data are normal, the points fall on the diagonal. It beats a histogram because the tails, precisely where the danger is, become two nearly invisible bars in a histogram and fill the ends of a QQ plot.

Reading the shapes. Ends escaping outward, a lying S, mean heavy tails: more extremes than a normal allows (multipath, faults, outliers). Ends curling inward mean light tails (a uniform, saturation). A banana means skew. Horizontal steps mean discrete data, quantization.

Shapiro–Wilk measures how straight the QQ line is: WWW near one is very normal. It is among the most powerful tests for small and medium samples. Anderson–Darling compares the cumulative distribution of the data with a normal's, with extra weight on the tails, which is exactly what matters for a filter, since the tails produce the disastrous corrections.

Do not obsess over the p-value. These tests answer whether the data are exactly normal, and in the real world the answer is always no. With 10,000 points they reject deviations too small to matter; with 20 they reject almost nothing. All of them assume independent samples, and with autocorrelated noise their p-value does not mean what it seems (autocorrelation).

Use them as an engineer: look at the QQ plot, ask whether the tails matter for the application, and treat the p-value as one more opinion, not the judge. What to do with each diagnosis is in sensor calibration.