mathematics//optimization//gradient descent//momentum

Momentum is a modification of gradient descent that keeps a running velocity, a decaying sum of past gradients, and steps along that velocity instead of along the latest gradient alone; it is part of almost every optimizer used to train neural networks (stochastic gradient descent with momentum, and the first moment of Adam) because it crosses long narrow valleys far faster than plain descent. The hiker feeling the slope in fog becomes a heavy ball rolling with friction.


Momentum is a modification of gradient descent that keeps a running velocity, a decaying sum of past gradients, and steps along that velocity instead of along the latest gradient alone; it is part of almost every optimizer used to train neural networks (stochastic gradient descent with momentum, and the first moment of Adam) because it crosses long narrow valleys far faster than plain descent. The hiker feeling the slope in fog becomes a heavy ball rolling with friction.

vk+1=β vk+∇L(θk),θk+1=θk−η vk+1v_{k+1} = \beta\,v_k + \nabla L(\theta_k),\qquad \theta_{k+1} = \theta_k - \eta\,v_{k+1}vk+1​=βvk​+∇L(θk​),θk+1​=θk​−ηvk+1​

vvv is the velocity and β\betaβ, typically 0.9, the fraction of it kept at each step. In a narrow valley the component of the gradient that bounces from wall to wall changes sign every step, so it cancels in the sum; the component along the valley floor keeps its sign, so it accumulates, up to 1/(1−β)=101/(1-\beta)=101/(1−β)=10 times a single gradient at β=0.9\beta=0.9β=0.9. The oscillation dies and the progress along the floor multiplies.

Momentum turns the optimizer into a damped second-order system, with β\betaβ setting how little friction there is. Too high a β\betaβ and the ball overshoots the minimum and swings around it, exactly as an underdamped mass on a spring does.

Tuning β\betaβ and η\etaη together matters. Since the velocity grows to about 1/(1−β)1/(1-\beta)1/(1−β) times the gradient, the effective step is about η/(1−β)\eta/(1-\beta)η/(1−β): raising β\betaβ from 0.9 to 0.99 multiplies it by ten, so the learning rate usually comes down when momentum goes up (learning rate).

With minibatch gradients the velocity also averages the noise of successive estimates, which is part of why stochastic gradient descent with momentum trains smoothly. Adam keeps the same running average as its first moment and adds a per-parameter scale from the second.

Its weakness shows on curved valleys. On Rosenbrock's banana valley in the figure of gradient descent, momentum runs along the floor but tends to overshoot at the bend, where plain descent crawls and Adam keeps steps of similar size in both directions; the momentum slider there shows the trade.