mathematics//numerical methods//finite differences
Finite differences are a family of formulas that approximate a derivative by evaluating a function at nearby points and dividing the change by the spacing, and they are the quickest way to get a derivative when the function is a black box: a simulator, a plant measured on the bench, a piece of code nobody wants to differentiate by hand. The forward difference is the definition of the derivative with the limit left out:
Finite differences are a family of formulas that approximate a derivative by evaluating a function at nearby points and dividing the change by the spacing, and they are the quickest way to get a derivative when the function is a black box: a simulator, a plant measured on the bench, a piece of code nobody wants to differentiate by hand. The forward difference is the definition of the derivative with the limit left out:
f′(x)≈f(x+h)−f(x)h,f′(x)≈f(x+h)−f(x−h)2h.f'(x)\approx\frac{f(x+h)-f(x)}{h},\qquad f'(x)\approx\frac{f(x+h)-f(x-h)}{2h}.f′(x)≈hf(x+h)−f(x),f′(x)≈2hf(x+h)−f(x−h).
The first costs one extra evaluation and has an error proportional to hhh; the second, the central difference, costs two and has an error proportional to h2h^2h2, so it is usually worth the extra call.
Choosing hhh is a trade between two errors that pull in opposite directions. A large step misses the curvature (truncation error); a small step subtracts two nearly equal numbers, and the rounding in each is divided by a tiny hhh (floating-point arithmetic). In double precision the best forward step is around 10−810^{-8}10−8 times the size of xxx, and the result keeps only about half of the sixteen digits; in single precision it keeps three or four. A step chosen without thinking can return a derivative that is pure rounding noise.
They are enough for simple functions and few parameters. Checking a hand-derived Jacobian for an extended Kalman filter, estimating the sensitivity of a simulated plant to one parameter, or linearizing a model around an operating point once, offline, are all jobs for a finite difference. Each input needs its own perturbed evaluation, so the cost grows with the number of inputs, and for a network with millions of weights automatic differentiation is the only practical choice.
On measured signals, the step is the sampling period and the function values carry sensor noise, so the difference amplifies that noise by 1/h1/h1/h. This is the naive derivative of a PID, and it is why the derivative term is filtered or taken from a sensor that measures rate directly (derivative action).
The same idea discretizes space. Replacing the spatial derivatives of a PDE by differences between neighbouring nodes turns heat flow along a bar into a large system of ODEs, which is how most simple thermal and diffusion models are simulated.