ML//continual learning
Continual learning is the family of methods that let a model keep acquiring knowledge or skills after deployment, from data that arrives over time, without losing what it already knew, and it is the missing piece between a model frozen at its training date and a system that gains experience the way a technician does. A deployed language model stops at its knowledge cutoff; a robot moved to a new site meets floors, lighting and parts its training never showed it.
Continual learning is the family of methods that let a model keep acquiring knowledge or skills after deployment, from data that arrives over time, without losing what it already knew, and it is the missing piece between a model frozen at its training date and a system that gains experience the way a technician does. A deployed language model stops at its knowledge cutoff; a robot moved to a new site meets floors, lighting and parts its training never showed it.
Take a delivery robot that trips on a step it had never seen. Three different things can happen next, and they are easy to confuse.
Inference only. A trained policy detects the obstacle and reacts. Nothing is kept: tomorrow it trips the same way and reacts the same way (inference).
External memory. The system writes down where the step is and avoids it on the next visit. Behaviour changes, the weights do not (agent memory).
Learning. The policy or its predictive model is updated from the experience, so it handles steps in general better, including ones it never met (online learning).
Storing what happened is not learning from it.
Memory changes what the system can look up; learning changes what it can do. Only the second generalizes to the next unseen step, and only the second can break skills that used to work.
The central obstacle is catastrophic forgetting: updating weights for the new data can degrade old abilities, because the same weights serve many tasks. Learning also has to resist bad data (one mislabelled episode, a faulty sensor) and must not introduce dangerous behaviour in a machine that moves.
The remedies come in four kinds: mixing old examples back into training (experience replay), penalizing changes to weights that mattered for old tasks (regularization such as elastic weight consolidation), training small add-on parameters while the base stays frozen (LoRA and adapters), and keeping knowledge outside the weights altogether (external memory, RAG).
Control engineering has learned online for decades. An adaptive controller re-estimates a handful of plant parameters in flight with recursive least squares, and it is accepted because the parameters are few, their meaning is physical and stability can be argued. A neural policy with millions of weights offers none of that yet, which is why most fielded systems learn offline and only remember online.