ML//kernel methods//Gaussian process

A Gaussian process is a kernel regression method that predicts a function from a few measured points and returns, at every input, both an estimate and its uncertainty; it is used where each measurement is expensive and knowing how sure the model is matters as much as the value. Feed it twenty bench runs of a drone motor (thrust against speed and temperature) and ask about a speed never tested: it answers with a mean and a variance, and the variance is small near the runs and grows between and beyond them.


A Gaussian process is a kernel regression method that predicts a function from a few measured points and returns, at every input, both an estimate and its uncertainty; it is used where each measurement is expensive and knowing how sure the model is matters as much as the value. Feed it twenty bench runs of a drone motor (thrust against speed and temperature) and ask about a speed never tested: it answers with a mean and a variance, and the variance is small near the runs and grows between and beyond them.

The idea is the kernel read as a statement about functions. The kernel k(x,x′)k(x,x')k(x,x′) says how strongly the values at two inputs are expected to move together: with the Gaussian kernel, two points closer than the length scale ℓ\ellℓ have nearly equal values, and two points far apart are unrelated. Conditioning on the measured points with Bayes' rule gives, at a new input x∗x_*x∗​,

μ(x∗)=k∗T(K+σn2I)−1y,σ2(x∗)=k(x∗,x∗)−k∗T(K+σn2I)−1k∗,\mu(x_*)=k_*^{\mathsf T}(K+\sigma_n^2 I)^{-1}y,\qquad \sigma^2(x_*)=k(x_*,x_*)-k_*^{\mathsf T}(K+\sigma_n^2 I)^{-1}k_*,μ(x∗​)=k∗T​(K+σn2​I)−1y,σ2(x∗​)=k(x∗​,x∗​)−k∗T​(K+σn2​I)−1k∗​,

where KKK holds the kernel between all training points, k∗k_*k∗​ between the new point and each of them, and σn2\sigma_n^2σn2​ is the measurement noise. The mean is a weighted sum of the measured values, much like kernel regression with weights fitted to the data; the variance starts at the prior and is reduced by every nearby measurement.

The variance is the product.

A model that knows where it has not looked can choose where to look next. That is the engine of Bayesian optimization, which tunes the gains of a controller on a test bench in a few dozen trials by measuring next where the predicted performance is good or the uncertainty is large.

The cost grows fast with data. Solving with KKK takes time cubic in the number of points NNN and memory quadratic, so a GP is comfortable with hundreds to a few thousand points and impractical past tens of thousands. On a large table, gradient boosting has replaced it; the GP keeps the niche of small, expensive datasets.

The kernel and its length scale are the model. They are fitted by maximizing the likelihood of the data, and a wrong choice shows as confident nonsense: a length scale too long smooths away a real bend, one too short makes the uncertainty collapse only on top of the points.

The variance is honest only inside the assumptions. It measures the distance to the data in kernel terms, not every way the world can differ; a sensor that drifted between runs is invisible to it (out-of-distribution inputs still need their own check).

Its relatives in the family are the support vector machine, which keeps only the boundary points, and k-nearest neighbors, which keeps all of them and no uncertainty.