mathematics//dynamical systems//impulse response
The impulse response of a linear time-invariant system is the output it produces, starting from rest, after a single unit kick at its input, and it describes the system completely: the output for any other input is the sum of shifted and scaled copies of it. Engineers meet it when they tap a structure with an instrumented hammer to find its modes, when they identify a process from data, and whenever they design a digital filter, whose coefficients are its impulse response. In discrete time the output is the input convolved with it,
The impulse response of a linear time-invariant system is the output it produces, starting from rest, after a single unit kick at its input, and it describes the system completely: the output for any other input is the sum of shifted and scaled copies of it. Engineers meet it when they tap a structure with an instrumented hammer to find its modes, when they identify a process from data, and whenever they design a digital filter, whose coefficients are its impulse response. In discrete time the output is the input convolved with it,
yk=∑j=0∞hj uk−j,y_k=\sum_{j=0}^{\infty}h_j\,u_{k-j},yk=j=0∑∞hjuk−j,
so hjh_jhj is how much an input applied jjj steps ago still weighs on the output now. A short impulse response is a short memory; one that decays slowly is a system that remembers.
The simplest case shows it. For xk+1=a xk+b ukx_{k+1}=a,x_k+b,u_kxk+1=axk+buk with yk=xky_k=x_kyk=xk, a unit kick produces hj=b aj−1h_j=b,a^{j-1}hj=baj−1 for j≥1j\ge1j≥1: a geometric decay if ∣a∣<1|a|<1∣a∣<1, and an input keeps mattering for roughly 1/(1−a)1/(1-a)1/(1−a) steps. For a general state-space model the sequence is
h0=D,hj=C Aj−1B(j≥1),h_0=D,\qquad h_j=C\,A^{j-1}B\quad (j\ge1),h0=D,hj=CAj−1B(j≥1),
the Markov parameters of an identification course: the input enters through BBB, is carried j−1j-1j−1 steps by AAA and read through CCC. The continuous version is the convolution integral of delay and lag.
It ties the classic descriptions together. The step response is its running sum, which is why plants are tested with steps (a step is easy to apply, a perfect kick is not) and differentiated afterwards; its Fourier transform is the frequency response, so the three views carry the same information for a linear system.
Learning it directly is a model choice. An FIR model estimates the hjh_jhj one by one from input and output records; it needs no structure, and needs as many coefficients as the memory is long, which for a slow thermal process sampled fast is hundreds (system identification). A state-space or ARX model packs the same decay into a few numbers.
Sequence models reuse the identity. An S4 layer is a linear state-space system whose output is its input convolved with its own Markov parameters, so it trains in parallel as a convolution and runs step by step as a recurrence (S4, convolution).
It exists only for linear time-invariant systems. A nonlinear plant answers a large kick differently from a small one, and a time-varying one differently at each instant, so a single sequence no longer describes it.