ML//SSM//S4

S4 (structured state space sequence model) is a neural network layer, introduced in 2021, that is a discretized linear state-space system with learned matrices and a specially structured state matrix \(\bar A\) designed to remember long histories, and it was the first model of the SSM family to handle sequences of many thousands of steps well. It trains like a convolution and runs like a recurrence, so it can learn from long sensor records in parallel on a GPU and then run sample by sample on modest hardware.


S4 (structured state space sequence model) is a neural network layer, introduced in 2021, that is a discretized linear state-space system with learned matrices and a specially structured state matrix Aˉ\bar AAˉ designed to remember long histories, and it was the first model of the SSM family to handle sequences of many thousands of steps well. It trains like a convolution and runs like a recurrence, so it can learn from long sensor records in parallel on a GPU and then run sample by sample on modest hardware.

sk+1=Aˉ sk+Bˉ uk,yk=C sk+D uks_{k+1}=\bar A\,s_k+\bar B\,u_k,\qquad y_k=C\,s_k+D\,u_ksk+1​=Aˉsk​+Bˉuk​,yk​=Csk​+Duk​

Because the layer is linear and time-invariant, its output is the input convolved with the impulse response (CBˉ, CAˉBˉ, CAˉ2Bˉ,… )(C\bar B,,C\bar A\bar B,,C\bar A^2\bar B,\dots)(CBˉ,CAˉBˉ,CAˉ2Bˉ,…), the same Markov parameters a system identification course estimates from a step test. During training the whole kernel is computed at once and applied as one long convolution; at inference the same layer is stepped as a recurrence with constant cost per sample. The structure of Aˉ\bar AAˉ is what makes that long kernel cheap to compute and keeps the memory stable over long horizons (its eigenvalues stay inside the unit circle, the discrete-time stability condition).

Between layers sit nonlinearities, so a stack of S4 layers is a nonlinear model even though each layer is linear.

Its limit is its virtue: the same Aˉ\bar AAˉ, Bˉ\bar BBˉ and CCC apply to every input, so it cannot decide to store one element and ignore the next. Mamba removes that limit by making them depend on the input, at the price of the convolution.

It is a good reminder for an engineer that a learned layer can be a textbook LTI system (state-space model): what is new is learning Aˉ\bar AAˉ, Bˉ\bar BBˉ, CCC end to end inside a network, while the dynamics are the ones in the textbook.