ML//kernel methods//support vector machine
A support vector machine is a classifier that separates two classes with the boundary leaving the widest possible margin to the closest examples of each, and it is used for classification with few examples and good features, where it often does well on a few hundred labelled cases that would starve a network. Classifying the operating mode of a compressor from six features and 300 examples is its kind of problem.
A support vector machine is a classifier that separates two classes with the boundary leaving the widest possible margin to the closest examples of each, and it is used for classification with few examples and good features, where it often does well on a few hundred labelled cases that would starve a network. Classifying the operating mode of a compressor from six features and 300 examples is its kind of problem.
Among all hyperplanes wTx+b=0w^{\mathsf T}x+b=0wTx+b=0 that separate the classes, it picks the one with the widest corridor to the nearest points. The margin measures 2/∥w∥2/|w|2/∥w∥, so widening it means shrinking ∥w∥|w|∥w∥, and points allowed inside the corridor pay a price:
minw,b 12∥w∥2+C∑i=1Nmax(0, 1−yi(wTxi+b)),yi∈{−1,+1}.\min_{w,b}\;\tfrac12\|w\|^2+C\sum_{i=1}^{N}\max\bigl(0,\,1-y_i(w^{\mathsf T}x_i+b)\bigr),\qquad y_i\in\{-1,+1\}.w,bmin21∥w∥2+Ci=1∑Nmax(0,1−yi(wTxi+b)),yi∈{−1,+1}.
The first term widens the corridor. The second, the hinge loss, is zero for a point on the right side with room to spare and grows linearly as a point invades the corridor or crosses it; CCC sets how much a violation costs against a narrower margin. The solution depends only on the points touching or crossing the margin, the support vectors: delete any other example and the boundary does not move.
The data enter only through inner products.
Written in its dual form, the classifier needs nothing but xiTxjx_i^{\mathsf T}x_jxiTxj between examples, so replacing that product by a kernel bends the boundary into curves without ever building the high-dimensional features. That kernel trick is what made the SVM the reference classifier of the late 1990s and 2000s.
It depends on the scales, as every method that measures distance does. A feature in thousands next to one in fractions dictates the margin alone; standardize first.
With a kernel it keeps a share of the training set as support vectors and evaluates the kernel against each of them, so its cost grows with the data: training goes from quadratic to cubic in NNN, and past tens of thousands of examples gradient boosting or a network usually wins on both accuracy and time. Shipping the support vectors to an edge device is the same burden as carrying a kNN history.
Its output is a signed distance to the boundary, not a probability. Turning it into one needs a post-hoc fit (probability calibration, where Platt scaling was born for exactly this), which matters as soon as a cost threshold is set on it.
The one-class SVM draws a boundary around normal data alone and flags what falls outside, one of the finer and costlier detectors of anomaly detection.
The nearest relative without the margin is logistic regression, also a hyperplane, trained on a smooth loss that never stops pushing.