ML//continual learning//experience replay
Experience replay is a training technique that keeps a store of past examples or interactions (the **replay buffer**) and mixes samples drawn from it into every update on new data, so that the gradient keeps seeing the old distribution while it learns the new one. It is used in two places: in reinforcement learning, where it made deep Q-learning stable, and in continual learning, under the name rehearsal, as the most reliable defence against catastrophic forgetting.
Experience replay is a training technique that keeps a store of past examples or interactions (the replay buffer) and mixes samples drawn from it into every update on new data, so that the gradient keeps seeing the old distribution while it learns the new one. It is used in two places: in reinforcement learning, where it made deep Q-learning stable, and in continual learning, under the name rehearsal, as the most reliable defence against catastrophic forgetting.
The reason it works is plain once the failure is seen. A network trained only on the latest stream follows the gradient of the latest stream, and nothing in that gradient protects the weights that old tasks relied on. Put a slice of old examples in each batch and the loss itself starts to object whenever an update would break them. In reinforcement learning the problem is correlation: consecutive steps of one episode are nearly identical, and learning from them in order makes the network chase its own last few seconds; sampling at random from a large buffer of transitions (st,at,rt,st+1)(s_t, a_t, r_t, s_{t+1})(st,at,rt,st+1) breaks that chain and lets each transition be reused many times (Q-learning is where this matters most).
It costs storage and raises a data question. Keeping old examples means keeping them legally and physically: a fleet of robots can retain logs from earlier sites, while a model trained on customer data may not be allowed to. When the data cannot be kept, generative replay trains a model to produce stand-ins for the old data, at the price of whatever that generator gets wrong.
The mix is a dial. Too little old data and forgetting returns; too much and the new task is learned slowly. Practitioners tune the ratio the way a forgetting factor is tuned in an estimator.
It sits beside three other remedies. Regularization (elastic weight consolidation) penalizes moving the weights that mattered for earlier tasks; adapters such as LoRA train new parameters and leave the old ones frozen; external memory keeps knowledge outside the weights entirely (agent memory). Replay is the simplest of the four to reason about, because what it protects is exactly what is in the buffer.
A robot retrained for a new warehouse that replays a share of its logs from the first one keeps its old competence by the same mechanism a student keeps algebra alive by doing an old problem now and then.