mathematics//similarity measure
A similarity measure is a function that takes two objects (two vectors, two sets, two signals, two strings) and returns a number saying how alike they are, with a **distance** being its mirror image that grows as they differ; it is the hidden decision inside every search engine, recommender, clustering, anomaly detector and duplicate finder, since all of them rank things by *closest*. Choosing the measure is choosing what counts as alike, and it is usually a more consequential decision than the algorithm built on top.
A similarity measure is a function that takes two objects (two vectors, two sets, two signals, two strings) and returns a number saying how alike they are, with a distance being its mirror image that grows as they differ; it is the hidden decision inside every search engine, recommender, clustering, anomaly detector and duplicate finder, since all of them rank things by closest. Choosing the measure is choosing what counts as alike, and it is usually a more consequential decision than the algorithm built on top.
The members of the family differ in the kind of object they compare and in what they ignore. Cosine similarity compares two vectors by their angle and ignores their length, which is right for embeddings and word counts, where a long document should not win just by being long. The Euclidean distance compares positions and keeps length, right for physical coordinates or a sensor's feature vector in consistent units. The Mahalanobis distance measures in units of the data's own spread and correlations, which is what a fault detector needs when temperature varies by tens of degrees and vibration by tenths. The Jaccard index compares sets by their overlap, for things that are present or absent (the tags on two documents, the alarms raised by two incidents, the components of two bills of materials).
The measure must match the representation.
Jaccard on embeddings or cosine on raw sets answers a question nobody asked; Euclidean distance on features with different units lets the largest number decide everything. Scaling the features, or picking a measure that is blind to the right things, comes before any threshold.
A true metric satisfies the triangle inequality (going through a third point is never shorter), which lets indexes prune a search; cosine similarity as such is not a metric, though related angular distances are, and many useful measures are not metrics at all.
In high dimensions distances concentrate: every point ends up roughly equally far from every other, and nearest neighbour loses meaning unless the data really live on a lower-dimensional structure (curse of dimensionality).
Attention inside a transformer is a learned similarity: a dot product between a query and keys, scaled and turned into weights (attention).
Similarity scores make poor evidence alone. A threshold on cosine for duplicate or on Jaccard for same incident is set from labelled pairs and checked with precision and recall, never guessed.