systems theory//engineering patterns//weighted averaging
Weighted averaging is the engineering pattern that reads almost every estimator as a mean of what was believed and what was just observed, each weighted by how reliable it is, and it is used to understand, compare and debug estimators by asking one question of each: where do its weights come from? A drone's altitude is the barometer's reading pulled toward the model's prediction, a fleet's agreed position is every agent's value pulled toward its neighbours', a transformer's output is every token's value pulled toward the ones most similar to the query. The arithmetic is the same; the source of the weights is what tells the methods apart.
Weighted averaging is the engineering pattern that reads almost every estimator as a mean of what was believed and what was just observed, each weighted by how reliable it is, and it is used to understand, compare and debug estimators by asking one question of each: where do its weights come from? A drone's altitude is the barometer's reading pulled toward the model's prediction, a fleet's agreed position is every agent's value pulled toward its neighbours', a transformer's output is every token's value pulled toward the ones most similar to the query. The arithmetic is the same; the source of the weights is what tells the methods apart.
In its simplest form there is a prior guess x^−\hat x^-x^−, a new observation zzz and one weight between them.
x^=(1−K) x^−+K z=x^−+K (z−x^−),0≤K≤1\hat x = (1-K)\,\hat x^- + K\,z = \hat x^- + K\,(z-\hat x^-),\qquad 0\le K\le 1x^=(1−K)x^−+Kz=x^−+K(z−x^−),0≤K≤1
The second form reads as a correction: start from the belief and move toward the observation by a fraction KKK of the surprise. With K=0K=0K=0 the observation is ignored, with K=1K=1K=1 the belief is thrown away, and every method in the family is a rule for choosing KKK (or one weight per source) at each step.
The weights carry the engineering. Two estimators with the same averaging line differ in what they know about trust: one computes it from measured noise and propagated uncertainty, another fixes it by hand, another learns it from data. Asking where the weights come from tells you what a method can be trusted to do when a sensor degrades.
In the Kalman filter the weight is the Kalman gain, K=P−/(P−+R)K=P^-/(P^-+R)K=P−/(P−+R): the model's propagated uncertainty against the sensor's noise variance. It recomputes the weight every step and returns its own remaining doubt, which is why it adapts when a GPS drops out. Fusing several sensors at once is inverse-variance weighting, weights proportional to 1/σi21/\sigma_i^21/σi2, the static core of sensor fusion; Bayes' rule gives the same answer for two Gaussian beliefs, since their product is precision-weighted.
In the complementary filter the weight comes from a time constant chosen by the designer: below a crossover frequency trust the accelerometer, above it the gyro. An EWMA fixes a single weight for all time. Recursive least squares derives its gain from an information matrix built from past data, discounted by a forgetting factor.
In the particle filter each hypothesis is weighted by the likelihood of the reading under it, so the weights come from the sensor model rather than from a variance. In the consensus protocol they come from the communication graph: each agent averages with its neighbours, and who hears whom decides the result.
In kernel regression and transformer attention the weight is a similarity: a kernel of distance, or a softmax of query-key dot products learned in training. This is the sharpest contrast in the family. Attention weights say how relevant a token is to the query, and nothing about how trustworthy it is; there is no noise model and no output uncertainty, so a confident, wrong source gets as much weight as a confident, right one.
Two witnesses describe the same car. The one who stood close and wore glasses gets more say than the one who glanced from across the street, and if both agree you are more certain than either alone. Every estimator in this note is that judge, with a different reason for trusting each witness.
The pattern stops where an average is the wrong summary: two well-separated hypotheses (a robot that is either in corridor A or corridor B) averaged into a position between them that is certainly wrong, which is the case the particle filter keeps as separate weighted points. The pattern's siblings are collected in engineering patterns; the dual question of how much a weight should react to surprise is residual as surprise.