ML//supervised learning//decision tree//random forest

A random forest is an ensemble of hundreds of deep decision trees, each trained on a random resample of the data and allowed to look at only a random subset of the features at every split, whose predictions are averaged; it is used as the first serious model on tabular data, because it works almost without tuning and sets the baseline every fancier model has to beat. Asked for the failure risk of a pump, three hundred trees each give an answer, and the forest reports their mean (or their vote).


A random forest is an ensemble of hundreds of deep decision trees, each trained on a random resample of the data and allowed to look at only a random subset of the features at every split, whose predictions are averaged; it is used as the first serious model on tabular data, because it works almost without tuning and sets the baseline every fancier model has to beat. Asked for the failure risk of a pump, three hundred trees each give an answer, and the forest reports their mean (or their vote).

Each tree on its own is a high-variance model that fits its sample in detail. Averaging cures variance only when the errors being averaged are not the same errors, so the forest decorrelates its trees twice. Bagging (bootstrap aggregating) trains each tree on a bootstrap resample of the rows, drawn with replacement, so every tree sees a slightly different dataset. The random feature subset at each split stops one dominant feature from making every tree start with the same question. With BBB trees whose errors have variance σ2\sigma^2σ2 and pairwise correlation ρ\rhoρ, the variance of the average is

ρ σ2+1−ρB σ2,\rho\,\sigma^2+\frac{1-\rho}{B}\,\sigma^2 ,ρσ2+B1−ρ​σ2,

so adding trees removes the second term and only lowering the correlation attacks the first, which is why the feature sampling matters as much as the resampling.

More trees never overfit; they only stop helping.

Past a few hundred, the average has converged and extra trees just cost time. The knobs that matter (features per split, minimum leaf size) have forgiving defaults, which is what makes the forest a baseline rather than a project.

It is robust where boosting is touchy. Without tuning it is rarely bad, while gradient boosting, which chains shallow trees each correcting the previous ones, usually ends a little more accurate once its learning rate and depth are tuned. A common order in practice is forest first, boosting when the extra accuracy pays.

It trains in minutes on a CPU for hundreds to millions of rows, the trees in parallel, and predicts in microseconds. The rows left out of each bootstrap give a free estimate of test error (the out-of-bag error), useful but no substitute for a split by time or by machine when the rows are correlated (train-validation-test split).

It shares every limit of trees. It cannot extrapolate beyond the training range, gives no usable gradient, and its average of three hundred trees is far harder to read than one shallow tree; feature importances help, and they are biased toward features with many possible split points.