ML//supervised learning
Supervised learning is the branch of machine learning that fits a function from examples whose correct output is known, the label, so that it predicts the output for new inputs; it is how a plant turns five years of sensor history and a spreadsheet of failures into an alarm, or a fleet turns battery tests into a pass or fail estimate. Each example pairs features \(x\) (temperatures, currents, hours of use) with a label \(y\), a number to predict (regression) or a class (classification), and learning means choosing, from a family of functions, the one whose predictions match the labels best without memorizing them. Without labels the problem is unsupervised learning; with three failures in five years it usually becomes anomaly detection, which learns what normal looks like instead.
Supervised learning is the branch of machine learning that fits a function from examples whose correct output is known, the label, so that it predicts the output for new inputs; it is how a plant turns five years of sensor history and a spreadsheet of failures into an alarm, or a fleet turns battery tests into a pass or fail estimate. Each example pairs features xxx (temperatures, currents, hours of use) with a label yyy, a number to predict (regression) or a class (classification), and learning means choosing, from a family of functions, the one whose predictions match the labels best without memorizing them. Without labels the problem is unsupervised learning; with three failures in five years it usually becomes anomaly detection, which learns what normal looks like instead.
Almost every method in the family is one line, empirical risk minimization: an average loss over the examples plus a penalty on complexity, minimized over a chosen family of functions. The members differ in the family. For a classifier the family fixes the shape of the decision boundary, the surface in feature space where the predicted class changes: a hyperplane for a logistic model, a staircase of axis-aligned cuts for a tree, curved regions for a kernel machine or a network.
Start at the top of the ladder and climb only when the data force it.
A physical rule with a threshold first, then a regression, then trees, then kernels or networks; each step costs data, maintenance and explainability, and must beat the one before by enough to pay for that (problem framing sets the baseline, applied ML holds the full table).
Linear regression predicts a number as a weighted sum of features, with weights that carry units and can be audited by eye. It is the baseline that never goes to waste, and it fits on a microcontroller.
Logistic regression passes the same sum through a sigmoid to give a probability. It is the auditable classifier, and a network with its hidden layers removed is exactly this.
A decision tree asks a chain of threshold questions that translate into nested conditions for a PLC. Alone it is unstable; as an ensemble (random forest, gradient boosting) it is the default for tabular data, with the caveat that trees never extrapolate.
Kernel methods learn by similarity to stored examples and suit small datasets, with uncertainty in the case of the Gaussian process. Data that are not a table (images, sound, raw signals, text) call for neural networks.
What the formula leaves out is the real objective, performance on data that do not exist yet; that is generalization, and measuring it honestly is evaluation.