ML//kernel methods//kernel regression
Kernel regression is a nonparametric regression method that predicts the value at a new input as a weighted average of all the measured values, each weighted by how similar its input is to the new one; it is used to interpolate a calibration map or a performance curve from scattered measurements without choosing a formula for it. The consumption of a pump at an operating point never measured is estimated from the points that were: the neighbours a few percent away in flow and head count heavily, the distant ones barely at all.
Kernel regression is a nonparametric regression method that predicts the value at a new input as a weighted average of all the measured values, each weighted by how similar its input is to the new one; it is used to interpolate a calibration map or a performance curve from scattered measurements without choosing a formula for it. The consumption of a pump at an operating point never measured is estimated from the points that were: the neighbours a few percent away in flow and head count heavily, the distant ones barely at all.
It is k-nearest neighbors made soft. Every example votes, with a weight given by a kernel k(x,xi)k(x,x_i)k(x,xi) that falls with distance, and the weights are normalized to add up to one (the Nadaraya-Watson estimator):
y^(x)=∑i=1Nk(x,xi)∑jk(x,xj) yi.\hat y(x)=\sum_{i=1}^{N}\frac{k(x,x_i)}{\sum_{j}k(x,x_j)}\,y_i .y^(x)=i=1∑N∑jk(x,xj)k(x,xi)yi.
The fraction is the share of attention example iii receives; the prediction is the mean of the yiy_iyi under those shares. With the Gaussian kernel the width ℓ\ellℓ sets how far the influence reaches, and choosing it is the whole modelling decision: a narrow kernel follows the noise from point to point, a wide one flattens real bends.
The same formula is attention.
Replace the fixed kernel by a learned similarity between a query and a set of keys, and the weighted mean of the values is the attention of a transformer. The step from 1964 statistics to modern networks is learning the kernel instead of choosing it.
A lookup table with linear interpolation, the calibration map in almost every engine controller and drive, is kernel regression with a triangular kernel on a grid. It runs on any microcontroller, which is why the elaborate kernels earn their place only with several input dimensions and scattered data.
Nothing is fitted, so nothing is extrapolated. Far from all examples the weights are all tiny and the normalization still forces them to sum to one, so the estimate collapses to the value of whichever example happens to be least far. Near the edges of the data the estimate is biased toward the interior.
The cost of a prediction grows with the dataset, as in every method that keeps its examples. The Gaussian process fits the weights to the data and adds the uncertainty this estimator lacks; it is the same idea with a probabilistic reading (weighted averaging collects the other places the pattern appears).