mathematics//calculus//Jacobian
The Jacobian is the matrix of all first partial derivatives of a vector function, one row per output and one column per input, and it is the local linear stand-in for a nonlinear map: whatever the function does far away, near a point it acts on small changes like this matrix. In a model \(\dot x=f(x,u)\), entry \((i,j)\) says how much the rate of state \(i\) changes when state \(j\) is moved a little:
The Jacobian is the matrix of all first partial derivatives of a vector function, one row per output and one column per input, and it is the local linear stand-in for a nonlinear map: whatever the function does far away, near a point it acts on small changes like this matrix. In a model x˙=f(x,u)\dot x=f(x,u)x˙=f(x,u), entry (i,j)(i,j)(i,j) says how much the rate of state iii changes when state jjj is moved a little:
Jij=∂fi∂xj.J_{ij}=\frac{\partial f_i}{\partial x_j}.Jij=∂xj∂fi.
For a pendulum with state angle and angular rate, θ˙=ω\dot\theta=\omegaθ˙=ω and ω˙=−(g/L)sinθ\dot\omega=-(g/L)\sin\thetaω˙=−(g/L)sinθ, the Jacobian has a 1 where the rate of the angle meets the angular rate, and −(g/L)cosθ-(g/L)\cos\theta−(g/L)cosθ where the angular acceleration meets the angle. Hanging (θ=0\theta=0θ=0) that entry is negative and the pendulum oscillates; inverted (θ=π\theta=\piθ=π) the cosine flips its sign and the same equation becomes unstable (pendulum).
A gradient and a Jacobian are both derivatives, of different shapes. The gradient belongs to a scalar function of many inputs (a loss, a cost) and has one entry per input; the Jacobian belongs to a function with several outputs and has one row per output. A gradient is the single row of the Jacobian of a scalar function, laid on its side. Confusing them is the usual source of transposed shapes in code.
It is the matrix every linear method actually uses on a nonlinear system. Linearization evaluates it at an equilibrium to get AAA and BBB; the extended Kalman filter re-evaluates it at every step around the current estimate; uncertainty propagation carries a covariance through it, Σy≈JΣxJ⊤\Sigma_y\approx J\Sigma_xJ^\topΣy≈JΣxJ⊤, which is how a sensor's error reaches metres of position in an error budget; Newton's method solves with it.
A neural network is a chain of Jacobians. Around the current input each layer acts as Jℓ=diag(ϕ′(zℓ)) WℓJ_\ell=\operatorname{diag}(\phi'(z_\ell)),W_\ellJℓ=diag(ϕ′(zℓ))Wℓ, and the whole network as their product; backpropagation multiplies them from the loss backwards. When their singular values sit below one the product shrinks layer after layer and the gradient vanishes (vanishing gradient).
It is valid at a point. A Jacobian taken at hover says nothing reliable about a 40° bank, and a Jacobian that loses rank at a point can declare a system uncontrollable or unobservable there although the nonlinear system is neither. In code it comes from a hand derivation, finite differences or automatic differentiation; a wrong hand-derived Jacobian is a classic silent bug in an EKF, and comparing it against finite differences is the cheap check.