ML//supervised learning//empirical risk minimization
Empirical risk minimization is the principle that trains a supervised model by choosing, from an allowed family of functions, the one with the smallest average loss on the training examples plus a penalty on its complexity; it is the one line that almost every method of supervised learning instantiates, from a linear fit of an oven's consumption to a network on a camera. The loss is what an error costs on one example, the average is the empirical risk, and the penalty keeps the winner from being a function that threads every noisy point.
Empirical risk minimization is the principle that trains a supervised model by choosing, from an allowed family of functions, the one with the smallest average loss on the training examples plus a penalty on its complexity; it is the one line that almost every method of supervised learning instantiates, from a linear fit of an oven's consumption to a network on a camera. The loss is what an error costs on one example, the average is the empirical risk, and the penalty keeps the winner from being a function that threads every noisy point.
f^=argminf∈F 1N∑i=1Nℓ(yi,f(xi))+λ Ω(f).\hat f=\arg\min_{f\in\mathcal F}\;\frac1N\sum_{i=1}^{N}\ell\bigl(y_i,f(x_i)\bigr)+\lambda\,\Omega(f).f^=argf∈FminN1i=1∑Nℓ(yi,f(xi))+λΩ(f).
xix_ixi are the features of example iii, yiy_iyi its label and NNN the number of examples. F\mathcal FF is the hypothesis class, the set of functions the learner may choose from (lines, trees of a given depth, networks of a given shape); ℓ\ellℓ is the loss function, Ω\OmegaΩ the complexity penalty and λ\lambdaλ its weight. Each method of the family is a choice of these three: squared loss over linear functions with no penalty is linear regression, the same with ∥w∥2|w|^2∥w∥2 is ridge regression, cross-entropy over a sigmoid of a linear score is logistic regression, hinge loss with ∥w∥2|w|^2∥w∥2 is the support vector machine.
What is wanted is not in the formula.
The objective is a small error on data that do not exist yet; the formula only sees the data at hand. Everything between the two (the penalty, the validation split, early stopping, more data) is the craft of generalization.
When the loss is minus the log-likelihood of the label under the model, ERM is maximum likelihood estimation with a flexible model, and the penalty plays the role of a prior: λ∥w∥2\lambda|w|^2λ∥w∥2 is a Gaussian prior on the weights, which makes the penalized fit a MAP estimation (regularization develops this).
The choice of F\mathcal FF is a modelling decision with consequences far beyond accuracy. A class of monotone functions, of piecewise-constant trees or of smooth networks fixes what the model can extrapolate, whether it can be read, and whether it gives a gradient an optimizer can use (inductive bias is the general name for these built-in assumptions).
Solving the minimization is optimization: closed form or QR for least squares, a convex problem with one valley for logistic regression, gradient descent with no guarantee for a network. A minimum of the empirical risk found exactly is still only as good as the examples are representative of the field (out-of-distribution inputs break that).