ML//time series//lag features
Lag features are columns built from past values of a time series (the value one step ago, one day ago, a rolling mean of the last hour), and they are used to turn a forecasting problem into an ordinary tabular regression that any model, usually gradient boosting, can learn. It is the method most industrial forecasts actually run on: predicting a plant's steam consumption for the next 24 hours from the production plan, the outside temperature and the last weeks of consumption.
Lag features are columns built from past values of a time series (the value one step ago, one day ago, a rolling mean of the last hour), and they are used to turn a forecasting problem into an ordinary tabular regression that any model, usually gradient boosting, can learn. It is the method most industrial forecasts actually run on: predicting a plant's steam consumption for the next 24 hours from the production plan, the outside temperature and the last weeks of consumption.
The construction is mechanical. Each row is one instant ttt. Its columns are lags of the target (yt−1y_{t-1}yt−1, yt−24y_{t-24}yt−24, yt−168y_{t-168}yt−168 for hourly data with daily and weekly rhythm), rolling statistics (a mean and a standard deviation over a window), calendar fields, and exogenous inputs known at decision time (the production plan for tomorrow, the weather forecast). The label is the value hhh steps later, yt+hy_{t+h}yt+h. Any regressor then learns the map from row to label, and the trees capture interactions nobody wrote down (high load matters more on cold days).
Every column must exist at the moment the forecast is made.
A rolling mean that includes the sample being predicted, or an exogenous value recorded after the fact (the actual temperature in place of the forecast one), gives a model that looks excellent in validation and fails on the first day of operation. That is data leakage, and in time series it is the default unless each feature is built with its timestamp in mind.
Against ARIMA, lag features win with many exogenous inputs and nonlinear effects, and lose with little data or a single smooth series, where ARIMA's few parameters generalize better. Both lose more often than expected to the persistence forecast at short horizons, which is why it is always reported beside them.
Several horizons are handled directly, one model per hhh, each trained on its own label; the alternative of feeding one-step predictions back as lags accumulates error with every step.
Validation splits by time, with a gap at least as long as the horizon between training and test, and never shuffles rows (cross-validation in its time-series form). Trees also cannot extrapolate: a load above anything in the training data gets the forecast of the highest load seen, a limit a physical feature or a linear term in the model can soften (physics-based features).