mathematics//probability//probability distribution

A probability distribution is the description of a random variable that says how much credibility each of its values carries, and it is what a model of noise, of failure times or of demand actually is once written down. For a variable with discrete values it is a list of probabilities that sum to one; for a continuous one it is a **density** \(p(x)\), whose area over an interval is the probability of landing there. A density is not itself a probability and can exceed one: a reading spread uniformly over half a millivolt has density 2 per millivolt. The **cumulative distribution function** \(F(x)=\Pr(X\le x)\) works for both and is what a percentile reads.


A probability distribution is the description of a random variable that says how much credibility each of its values carries, and it is what a model of noise, of failure times or of demand actually is once written down. For a variable with discrete values it is a list of probabilities that sum to one; for a continuous one it is a density p(x)p(x)p(x), whose area over an interval is the probability of landing there. A density is not itself a probability and can exceed one: a reading spread uniformly over half a millivolt has density 2 per millivolt. The cumulative distribution function F(x)=Pr⁡(X≤x)F(x)=\Pr(X\le x)F(x)=Pr(X≤x) works for both and is what a percentile reads.

A family is chosen from the mechanism that produced the quantity, because each family is a claim about how the numbers came to be.

A sum of many small independent effects gives the normal distribution (central limit theorem): the error of a machined dimension adds temperature, wear, play and vibration. It is the usual default, reasonable in the centre and often optimistic in the tails.

A product of effects gives the lognormal distribution, skewed to the right, used for repair times and particle sizes: what multiplies adds in the logarithm, and the logarithm becomes normal.

Maxima and minima follow extreme value distributions: the day's strongest gust, the weakest link of a chain, the life of the first bearing to fail in a fleet, which is where the Weibull distribution of reliability comes from.

A bounded quantity with no preferred value gives the uniform distribution: the rounding error of an ADC spreads evenly over one step Δ\DeltaΔ, with variance Δ2/12\Delta^2/12Δ2/12 (quantization).

Waiting times between independent events give the exponential distribution; few dominant causes, cascades and congestion give heavy tails, where rare huge values dominate the totals.

Estimated quantities have distributions of their own. Student's t distribution widens an interval when σ\sigmaσ itself was estimated from few data (with 25 readings the 95 % factor is 2.06 instead of 1.96, confidence interval), and the chi-square distribution governs sums of squared normal errors, which is how a filter is checked for honesty.

Choosing wrong rarely shows in the middle of the data, where most families look alike; it shows in the tails, which is where alarms, margins and failure probabilities live. A histogram with a few hundred samples cannot tell a normal from a heavy tail at the one-in-a-million level, so that choice has to come from the physics.