ML//imitation learning

Imitation learning is a way of obtaining a control policy by learning from demonstrations, recordings of an expert's states and actions, instead of from a reward or a model; it is used when showing the right behaviour is easier than specifying it, as with a pilot's landings logged for a drone, an operator's moves on a process unit, or teleoperated grasps for a robot arm. Its simplest form, **behavior cloning**, is plain supervised learning: a small network is trained to map each recorded state to the action the expert took, and then runs as the controller, often a few tens of microseconds per step on the vehicle's own processor.


Imitation learning is a way of obtaining a control policy by learning from demonstrations, recordings of an expert's states and actions, instead of from a reward or a model; it is used when showing the right behaviour is easier than specifying it, as with a pilot's landings logged for a drone, an operator's moves on a process unit, or teleoperated grasps for a robot arm. Its simplest form, behavior cloning, is plain supervised learning: a small network is trained to map each recorded state to the action the expert took, and then runs as the controller, often a few tens of microseconds per step on the vehicle's own processor.

It is cheap where the alternatives are dear. Reinforcement learning needs a reward that cannot be gamed and millions of simulated steps; a classical design needs a model. A few hours of demonstrations need neither, and they encode tricks an expert never wrote down.

In closed loop, a small error compounds.

The policy's own outputs decide which states it will see next: a slight mistake carries the vehicle to a state the expert never visited, where the network is worse, so it errs more and drifts further. The training data came from the expert's trajectories, and the policy changes the distribution it is tested on.

The growth is real and measurable: in the analysis that made the problem famous, the expected cost of cloning grows with the square of the horizon where an ordinary classifier's error grows linearly. The standard remedy collects expert corrections on the states the learned policy actually reaches and retrains on them (the DAgger scheme), at the price of keeping an expert in the loop.

The same mechanism appears in language models, where a model trained only on human text drifts once it feeds on its own outputs (exposure bias), and it is a special case of out-of-distribution inputs, here produced by the controller itself.

A learned policy can also have large local slopes, a high gain inside the loop, so in critical systems it is wrapped in monitoring of its inputs, limits on its outputs and a verified classical controller ready to take over (safety filter). It is one member of learning-based control, often a first stage before reinforcement learning refines the policy in simulation.