ML//RL//temporal-difference error

The temporal-difference error is the gap between what the Bellman equation says a value should be, given the reward just received and the value of the state just reached, and what the agent's current estimate says it is; it is the learning signal that drives Q-learning and most of reinforcement learning. For action values it reads


The temporal-difference error is the gap between what the Bellman equation says a value should be, given the reward just received and the value of the state just reached, and what the agent's current estimate says it is; it is the learning signal that drives Q-learning and most of reinforcement learning. For action values it reads

δ=r+γmax⁡a′Q(s′,a′)−Q(s,a),\delta = r + \gamma \max_{a'} Q(s',a') - Q(s,a),δ=r+γa′max​Q(s′,a′)−Q(s,a),

the bracket of the Q-learning update. The first two terms are a fresh, one-step guess built from experience; the last is the old belief. A positive δ\deltaδ means the step went better than expected and the estimate is raised by αδ\alpha\deltaαδ; a negative one lowers it.

Take a delivery drone whose estimate says that leaving the depot with this battery level is worth 10. It flies one leg, earns 1, and lands in a state its table values at 8, with γ=0.9\gamma=0.9γ=0.9. The one-step guess is 1+0.9⋅8=8.21+0.9\cdot 8=8.21+0.9⋅8=8.2, so δ=−1.8\delta=-1.8δ=−1.8: the depot start was worth less than it thought. Nothing waited for the end of the mission; the correction used the estimate of the next state, which is called bootstrapping, and it is what lets the agent learn during a task that never ends.

The TD error is a residual.

It is expected minus observed, the same kind of surprise signal as the innovation that corrects a Kalman filter and the residual that raises a fault alarm (residual as surprise).

Its average shrinks to zero only when the value function satisfies Bellman. A TD error that stays large in some states says the estimate there is still wrong, or that the world changed; one that never settles anywhere points to a learning rate too high or a reward that is noisy.

Bootstrapping is also its weakness. The target leans on the agent's own estimate, so errors propagate from state to state, and with a neural network approximating QQQ this feedback can diverge, which deep RL tames with tricks such as a slowly updated copy of the network for the target.