ML//time series//autoregressive model
An autoregressive model is a linear time-series model that predicts the current value of a series as a weighted sum of its \(p\) previous values plus an unpredictable noise term, and it is the first real model to try after the persistence forecast on slow process variables, demand curves or temperatures. Written as AR(\(p\)):
An autoregressive model is a linear time-series model that predicts the current value of a series as a weighted sum of its ppp previous values plus an unpredictable noise term, and it is the first real model to try after the persistence forecast on slow process variables, demand curves or temperatures. Written as AR(ppp):
yt=c+∑i=1pϕi yt−i+εty_t = c + \sum_{i=1}^{p}\phi_i\, y_{t-i} + \varepsilon_tyt=c+i=1∑pϕiyt−i+εt
Each coefficient ϕi\phi_iϕi says how strongly the value iii steps back pulls the present, ccc is a constant and εt\varepsilon_tεt is white noise, the part no amount of history predicts. The equation is linear in the coefficients, so fitting is least squares on a table of lagged values: no iterative training, an answer in milliseconds on a laptop, and coefficients a reader can interpret (a ϕ1\phi_1ϕ1 close to 1 is a slow process with a long memory).
Stacking the past values into a vector turns the model into a state-space model, the companion form. For AR(2):
[ytyt−1]=[ϕ1ϕ210][yt−1yt−2]+[10]εt\begin{bmatrix} y_t\\ y_{t-1}\end{bmatrix}=\begin{bmatrix}\phi_1&\phi_2\\ 1&0\end{bmatrix}\begin{bmatrix} y_{t-1}\\ y_{t-2}\end{bmatrix}+\begin{bmatrix}1\\ 0\end{bmatrix}\varepsilon_t[ytyt−1]=[ϕ11ϕ20][yt−1yt−2]+[10]εt
The state is the last two values, and the matrix carries the whole behaviour of the series. That is what lets every tool written for state-space models (a Kalman filter, a stability test, a simulator) accept an AR model unchanged.
The series settles around a fixed mean, and its forecasts decay toward it, exactly when every eigenvalue of that matrix lies inside the unit circle, the usual discrete-time stability test. A random walk is the AR(1) with ϕ1=1\phi_1=1ϕ1=1, an eigenvalue sitting on the circle: shocks never fade, which is why such a series is differenced before modelling (ARIMA).
With an input added (a valve opening, a heater command) the same regression becomes the ARX model, the first rung of system identification (ARX model).
The name collides with autoregressive generation in language models, where each token is sampled conditioned on the ones before it. Both feed their own outputs back; an AR(ppp) is a linear model with a handful of coefficients, a language model a network with billions.
Its states are past values, a black box. A model whose state is a temperature and a heat balance extrapolates to conditions never recorded; an AR model only continues the patterns it saw, which makes it an honest, cheap starting point when the physics is unknown and a poor one when it is known.