ML//neural network//multilayer perceptron
A multilayer perceptron (**MLP**) is a neural network made of fully connected layers in series, each one a weighted sum of all the outputs of the previous layer passed through a nonlinearity, and it is the default model for a fixed-length vector of inputs: forty features of a pump, the state of a drone fed to a learned policy, the per-token block inside a transformer. Every unit of a layer sees every unit of the one before, so it assumes nothing about the structure of its input, which is both its generality and why it needs more data than a CNN on an image.
A multilayer perceptron (MLP) is a neural network made of fully connected layers in series, each one a weighted sum of all the outputs of the previous layer passed through a nonlinearity, and it is the default model for a fixed-length vector of inputs: forty features of a pump, the state of a drone fed to a learned policy, the per-token block inside a transformer. Every unit of a layer sees every unit of the one before, so it assumes nothing about the structure of its input, which is both its generality and why it needs more data than a CNN on an image.
h(ℓ)=ϕ(W(ℓ)h(ℓ−1)+b(ℓ)),h(0)=x,ℓ=1,…,Lh^{(\ell)}=\phi\left(W^{(\ell)}h^{(\ell-1)}+b^{(\ell)}\right),\qquad h^{(0)}=x,\quad \ell=1,\dots,Lh(ℓ)=ϕ(W(ℓ)h(ℓ−1)+b(ℓ)),h(0)=x,ℓ=1,…,L
Here h(ℓ)h^{(\ell)}h(ℓ) is the vector of outputs of layer ℓ\ellℓ (its activations), W(ℓ)W^{(\ell)}W(ℓ) a weight matrix, b(ℓ)b^{(\ell)}b(ℓ) a bias vector and ϕ\phiϕ the activation function, applied element by element: a ReLU or a tanh. The last layer is usually linear for regression, or a sigmoid for a yes or no answer. A layer from 64 to 64 units has 64⋅64+64=416064\cdot64+64=416064⋅64+64=4160 parameters, so a network with 40 inputs, two hidden layers of 64 and one output has 6,849.
accuracy68 % loss0.556 parameters9 epochs60 A network with 1 hidden layer of 2 ReLU neurons (9 parameters) learns to separate two rings. After 60 epochs it classifies 68 % of the training points correctly, with cross-entropy 0.556.
On the rings, one hidden layer of 2 neurons cannot close the boundary (two half-planes enclose nothing); raise it to 3 or 4 and a polygon appears. Then try XOR and the spiral, which ask for more layers, compare ReLU's straight edges with tanh's curves, and push the learning rate up until the loss starts to jump.
Remove the hidden layers and you have logistic regression. A hidden layer of nnn ReLU units cuts the plane with nnn lines and the output combines the pieces, so the network is piecewise linear; tanh units give smooth curves instead.
Width and depth trade differently. One wide hidden layer can already approximate any continuous function on a bounded region (universal approximation theorem); depth reuses features and reaches the same function with far fewer parameters, at the price of harder training (vanishing gradient).
It is cheap and predictable at run time. Two hidden layers of 64 are about 5,000 multiply-accumulates, tens of microseconds on a fast Cortex-M, compatible with a 1 kHz loop, and the same operations every time, so its worst case is measured once (TinyML).
On tables it rarely wins. With hundreds or thousands of rows of process data, gradient boosting usually matches or beats it; an MLP earns its place as a part of larger networks (the feed-forward network of each transformer block) and as a small learned controller (learning-based control).