ML//supervised learning//decision tree

A decision tree is a supervised model that predicts by asking a chain of threshold questions about the features, each answer sending the case down one branch until a leaf gives the prediction; it is used where a rule must be readable and cheap to run, and as the building block of the ensembles that dominate tabular data. *Is the bearing temperature above 72 °C? Then, is the vibration above 7 mm/s?* Each question splits the data in two, and each leaf returns the mean of its training examples (regression) or the share of failures among them (classification).


A decision tree is a supervised model that predicts by asking a chain of threshold questions about the features, each answer sending the case down one branch until a leaf gives the prediction; it is used where a rule must be readable and cheap to run, and as the building block of the ensembles that dominate tabular data. Is the bearing temperature above 72 °C? Then, is the vibration above 7 mm/s? Each question splits the data in two, and each leaf returns the mean of its training examples (regression) or the share of failures among them (classification).

The tree is grown greedily. At each node the algorithm tries every feature and every threshold and keeps the split that most reduces the impurity of the two halves, typically the Gini impurity G=1−∑cpc2G=1-\sum_c p_c^2G=1−∑c​pc2​, with pcp_cpc​ the fraction of class ccc in the node: zero when the node is pure, highest when the classes are evenly mixed. A shallow tree reads at a glance and translates into nested IF statements that run on any PLC in a few microseconds.

One tree has enormous variance.

Change a few training rows and the first split can change, and with it the whole tree below. The two ensembles exist to cure this: averaging many trees, or chaining many small ones.

Trees do not extrapolate. A tree is a piecewise-constant function, so outside the training range it returns the last step: if the bearing never ran above 70 °C in the data, at 90 °C the model answers as it did at 70 °C. A model with physics, even a linear one, degrades more gracefully when the plant moves to conditions it has not seen.

Their gradients are useless. Flat almost everywhere and jumping at the thresholds, a tree gives an optimizer nothing to follow, which rules it out as the internal model of an MPC that needs derivatives.

What they do well is tabular data with mixed scales and gaps: thresholds need no scaling, irrelevant columns are simply never chosen, interactions between features are captured without being written, and missing values can be routed down a branch.

The two members answer the variance in opposite ways. A random forest trains hundreds of deep trees independently on resampled data and averages them; it works almost without tuning and is the baseline to beat. Gradient boosting chains hundreds of shallow trees, each fitted to what the previous ones got wrong; it squeezes out the most accuracy and is the industry default for tables, at the price of more hyperparameters to tune.