ML//evaluation//train-validation-test split
A train-validation-test split is the division of a dataset into three disjoint parts with three jobs: the training set fits the model's parameters, the validation set chooses among models and their **hyperparameters** (the settings fixed outside the fit, such as a regularization weight \(\lambda\), a tree depth or a learning rate), and the test set is looked at once, at the end, to estimate the error on new data. If the model is chosen by looking at the test set, the test set has become a second validation set and no longer estimates anything.
A train-validation-test split is the division of a dataset into three disjoint parts with three jobs: the training set fits the model's parameters, the validation set chooses among models and their hyperparameters (the settings fixed outside the fit, such as a regularization weight λ\lambdaλ, a tree depth or a learning rate), and the test set is looked at once, at the end, to estimate the error on new data. If the model is chosen by looking at the test set, the test set has become a second validation set and no longer estimates anything.
How the rows are divided decides which question the final number answers, and on industrial data the default random shuffle answers none of the useful ones. Sensors sampled every second produce consecutive rows that are near twins, so a shuffle puts twins on both sides and the test measures memory (data leakage).
The split is the question.
Split by time, with a gap, to ask whether the model works in the future of these machines; by group, to ask whether it works on a machine it has never seen; by site, to ask whether it transfers to another plant. Never shuffle a series.
The temporal split with gap trains on the past and tests on the future, leaving between them a gap as long as the feature windows plus the prediction horizon, so that no window straddles the boundary. With one row per minute and day-long windows, the gap is at least 1,440 rows (TimeSeriesSplit with gap=1440 in scikit-learn). Shuffling a series is sitting the exam with the next day's answers stapled to the back.
The group split puts all rows of a machine on the same side (GroupKFold). Otherwise the model learns to recognize machines, by their individual vibration fingerprints, instead of failures, and the score collapses on the first new unit.
Leave-one-plant-out holds out an entire site, or a manufacturing batch, to test transfer; it is the honest test before a model trained in one factory is installed in the next.
With few data the three parts are too small, and cross-validation rotates them; every rotation must keep the same grouping and time order, or it brings back the leakage the split was meant to prevent.