ML//RL//policy//policy gradient
Policy gradient methods are the family of reinforcement learning algorithms that adjust the parameters of a policy directly, by gradient descent on the expected reward, instead of first learning the value of every action and then picking the best one as Q-learning does. They are what trains continuous controllers in simulation (the walking policies of quadrupeds, a drone's racing policy) and, through PPO, the language models tuned with RLHF. Their advantage is that the action can be continuous (a motor torque, a thrust) or a whole stochastic distribution, where taking a maximum over actions at every step would be impossible.
Policy gradient methods are the family of reinforcement learning algorithms that adjust the parameters of a policy directly, by gradient descent on the expected reward, instead of first learning the value of every action and then picking the best one as Q-learning does. They are what trains continuous controllers in simulation (the walking policies of quadrupeds, a drone's racing policy) and, through PPO, the language models tuned with RLHF. Their advantage is that the action can be continuous (a motor torque, a thrust) or a whole stochastic distribution, where taking a maximum over actions at every step would be impossible.
The idea fits in a sentence: run the policy, see which episodes went well, and make the actions taken in them more likely. Written out, the gradient of the expected return with respect to the policy's parameters θ\thetaθ is
∇θJ=E[∇θlogπθ(at∣st) At],\nabla_\theta J=\mathbb E\Big[\nabla_\theta\log\pi_\theta(a_t\mid s_t)\,A_t\Big],∇θJ=E[∇θlogπθ(at∣st)At],
where ∇θlogπθ\nabla_\theta\log\pi_\theta∇θlogπθ is the direction that makes the chosen action more probable and AtA_tAt (the advantage) says how much better that action turned out than what was expected from that state. The plainest version, REINFORCE, uses the whole return of the episode for AtA_tAt, which works and is very noisy: one lucky gust that kept a drone up credits every action of the flight.
Actor-critic splits the job in two.
The actor is the policy; the critic is a learned value function that predicts the return from each state, so the advantage becomes the surprise relative to that prediction. Subtracting the expectation leaves the gradient's direction unchanged on average and cuts its variance enormously, which is what made these methods practical; PPO is an actor-critic with a limit on how far each update may move the policy.
They are hungry for samples. A gradient estimated from episodes needs many of them, so learning happens in a fast simulator with thousands of robots in parallel and the result is carried across the sim-to-real gap with domain randomization; a real plant could never supply that much trying.
Most of them are on-policy: the gradient is valid only for data produced by the current policy, so old episodes are discarded after each update, one more reason they are expensive.
The policy comes out as a network that guarantees nothing about stability or limits. In control it runs behind a safety layer and a classical fallback, as learning-based control describes.