ML//continual learning//online learning
Online learning is a training regime in which a model updates its parameters as each new datum or small batch arrives, instead of being fitted once on a fixed dataset (batch or offline learning), and it is used where the world drifts, where the stream is too large to store, or where each user or machine needs its own adjustment. The model is never finished: every observation is both a prediction to make and a lesson to absorb.
Online learning is a training regime in which a model updates its parameters as each new datum or small batch arrives, instead of being fitted once on a fixed dataset (batch or offline learning), and it is used where the world drifts, where the stream is too large to store, or where each user or machine needs its own adjustment. The model is never finished: every observation is both a prediction to make and a lesson to absorb.
The simplest case is an engineering one. An estimator on a pump that tracks the motor's torque constant with recursive least squares updates its estimate at every sample, and a forgetting factor sets how fast old samples lose weight: close to one, the estimate is steady and slow to follow a real change; lower, it follows wear quickly and jitters with the noise. Every online learner faces that same dial, called the stability-plasticity trade-off in machine learning.
It answers data drift. A spam filter, a demand forecast or a vibration model whose inputs change with season, wear or behaviour keeps its accuracy only if it keeps learning; retraining from scratch every night is the offline alternative, often simpler and easier to audit.
It learns its mistakes as fast as its lessons. A stuck sensor, a mislabelled event or a deliberate poisoning enters the weights immediately, with no human review between data and model, so production online learners carry guards: bounds on each update, outlier rejection, a frozen copy to fall back to.
In large neural networks it is rare. Each update costs a backward pass, can erode old skills (catastrophic forgetting) and makes the model's behaviour depend on the order of its inputs, which complicates testing and certification. Deployed language models therefore keep their weights fixed and adapt through context and memory.
Two near names mean other things. Online inference is a model answering live requests with fixed weights; writing facts to a database the agent can query (agent memory) changes behaviour without any parameter update. Neither is online learning.
The family it belongs to, and the remedies for forgetting, are in continual learning.