ML//time series//ARIMA

ARIMA is a family of linear time-series models that combines an autoregressive model with a moving average of past noise terms and a number of differencings, and it is the standard statistical forecaster for a single series that drifts: monthly energy consumption, a slowly fouling heat exchanger's efficiency, demand for a spare part. The family was set out by Box and Jenkins in 1970 and is written ARIMA(\(p\), \(d\), \(q\)).


ARIMA is a family of linear time-series models that combines an autoregressive model with a moving average of past noise terms and a number of differencings, and it is the standard statistical forecaster for a single series that drifts: monthly energy consumption, a slowly fouling heat exchanger's efficiency, demand for a spare part. The family was set out by Box and Jenkins in 1970 and is written ARIMA(ppp, ddd, qqq).

Each letter answers one problem. The AR part, ppp past values, captures how the series remembers itself. The MA part, qqq past noise terms, captures shocks that echo for a few steps and then stop (a one-off disturbance whose effect lasts two samples), something a short AR model would need many coefficients to imitate; together they make ARMA. The I part handles drift. A series that wanders like a random walk has no mean to return to, so it is first transformed by differencing,

Δyt=yt−yt−1,\Delta y_t = y_t - y_{t-1},Δyt​=yt​−yt−1​,

applied ddd times, and the ARMA model is fitted to the differences. One differencing removes a wandering level, two remove a wandering slope; the forecast is then summed back up.

Choosing ppp, ddd and qqq used to be read from autocorrelation plots; today libraries search a small grid and keep the model with the best information criterion. The seasonal extension (SARIMA, the S) adds the same three terms at a lag of one season, and exogenous regressors (the X of SARIMAX) let the production plan or the outside temperature enter as inputs, which brings it close to the ARX models of system identification.

Every ARIMA is a state-space model, and the libraries fit it that way. Missing samples, common in historian data, are handled without imputation, because a step with no measurement is simply a step of prediction.

It is fast and interpretable, and it is beaten in two places: by lag features with gradient boosting when many exogenous inputs act nonlinearly, and by the persistence forecast more often than its users expect at short horizons. Over-differencing is the classic error: differencing a series that was already stationary adds noise and makes the forecast worse.