ML//inductive bias
An inductive bias is an assumption about the world built into a learning method, which makes it prefer some explanations of the data over others before it has seen any; in neural networks it lives mostly in the architecture, and choosing an architecture is choosing which assumption to pay for. Every learner needs one. A finite set of examples is compatible with infinitely many functions, and the bias decides which of them the model picks in the gaps between the training points.
An inductive bias is an assumption about the world built into a learning method, which makes it prefer some explanations of the data over others before it has seen any; in neural networks it lives mostly in the architecture, and choosing an architecture is choosing which assumption to pay for. Every learner needs one. A finite set of examples is compatible with infinitely many functions, and the bias decides which of them the model picks in the gaps between the training points.
The trade shows in the data bill. A camera on an inspection line sees cracks anywhere in the frame. A CNN assumes that a pattern means the same wherever it appears, so a crack learned in one corner is recognized in all the others and a few hundred images suffice; a multilayer perceptron on the raw pixels assumes nothing, and would have to see cracks at every position to learn the same thing.
Each architecture is an assumption.
The CNN assumes translation, the RNN and the state-space models a state that evolves in time, the GNN that relations are local, and attention that any element may matter to any other. When the assumption holds it replaces data; when it is wrong it is an error that more data only partly removes.
CNN: translation equivariance, right in photographs and wrong in a spectrogram, where a pattern moved from 100 Hz to 200 Hz is another fault. Normalizing frequency by shaft speed restores the assumption instead of fighting it.
RNN and SSM: a fixed-size summary of the past, updated one step at a time, the same structure as a state-space model. It suits dynamics and virtual sensors, and it costs sequential computation and a long training through time.
GNN: influence travels along the edges of a known graph, so the same weights serve 10 drones or 200; it fails when the graph is wrong or the effect is long-range.
Attention: the weakest assumption of the four, any element may consult any other, which is why a transformer needs the most data (or a pretrained model) and pays a cost quadratic in length.
The idea is older than networks. A linear model assumes linearity, ridge regression prefers small weights (a prior, said in Bayesian terms), k-nearest neighbors assumes nearby points behave alike, and data augmentation writes invariances into the data. The strongest bias an engineer can supply is physics itself (physics-based features, grey-box model), and it is also the one that extrapolates.