ML//time series//persistence forecast

The persistence forecast is the baseline forecasting rule that predicts the future value of a series as its present value, \(\hat y_{t+h}=y_t\), and it is used as the yardstick every other forecaster has to beat before it deserves production. Its seasonal cousin, the **seasonal naive forecast**, repeats the value one season ago: the load at this hour yesterday, or on the same weekday last week.


The persistence forecast is the baseline forecasting rule that predicts the future value of a series as its present value, y^t+h=yt\hat y_{t+h}=y_ty^​t+h​=yt​, and it is used as the yardstick every other forecaster has to beat before it deserves production. Its seasonal cousin, the seasonal naive forecast, repeats the value one season ago: the load at this hour yesterday, or on the same weekday last week.

It is hard to beat for a physical reason. Most quantities an engineer forecasts are outputs of slow processes (a tank level, a building's temperature, a grid's load), and over a short horizon a slow process has barely moved. At a horizon of a few samples the persistence error is small, so a model that looks accurate in absolute terms may be adding nothing; the honest score is the error relative to persistence at the same horizon (the skill score, one minus the ratio of the two errors).

A model that does not beat persistence does not deserve production.

Report every forecaster beside persistence and the seasonal naive forecast, at the horizon the decision needs; many sophisticated models fail this test, and the ones that pass it show by how much.

Persistence also hides inside models. Trained to predict one step ahead at a high sampling rate, a network can learn to copy its last input and score a tiny error without having understood any dynamics; validating an identified model by free-run simulation, where its own predictions feed the next step, exposes the copy (model validation).

Its error grows with the horizon at a rate set by how fast the process forgets its present: a series with strong autocorrelation at lag hhh is well served by persistence at that horizon, and the autocorrelation function is the quickest estimate of where better models start to pay.

The rest of the ladder, from the autoregressive model and ARIMA to lag features with boosting, is laid out in time series; persistence is the same choice as the flyswatter rule made in forecasting.