ML//unsupervised learning//t-SNE
t-SNE (t-distributed stochastic neighbor embedding) is a nonlinear dimensionality reduction method that places high-dimensional points on a plane so that points that were near neighbours stay near, and it is used to look at data a human cannot see: hundreds of word embeddings, the internal features of a network, the vibration signatures of a fleet of bearings. It turns distances into probabilities of being neighbours (a Gaussian in the original space, a heavier-tailed Student t on the plane) and moves the points on the plane until the two sets of probabilities match.
t-SNE (t-distributed stochastic neighbor embedding) is a nonlinear dimensionality reduction method that places high-dimensional points on a plane so that points that were near neighbours stay near, and it is used to look at data a human cannot see: hundreds of word embeddings, the internal features of a network, the vibration signatures of a fleet of bearings. It turns distances into probabilities of being neighbours (a Gaussian in the original space, a heavier-tailed Student t on the plane) and moves the points on the plane until the two sets of probabilities match.
The map it draws is honest about one thing and only one: who is next to whom. Clusters that appear are usually real groups, which is why it became the standard picture of an embedding space, with pump, valve and compressor huddled in one corner and the shift names in another. Everything else on the picture is unreliable: the size of a cluster, the empty space between two clusters and their relative positions are artefacts of the method and of its main knob, the perplexity (roughly how many neighbours each point tries to keep), which here means something different from the language-model metric Perplexity.
t-SNE visualizes vectors; it never produces them.
The embeddings come from a trained model (Word2Vec, a transformer), and t-SNE only draws a flattened picture of them afterwards. The picture cannot be fed to another model, and two runs with different seeds draw two different maps of the same data.
Distances between clusters mean little. Two groups drawn far apart may be close in the original space; reading the gap as a degree of difference is the classic mistake.
PCA is the linear, deterministic alternative: it keeps large-scale distances and directions, can project new points, and runs in seconds, but it smears curved structure. UMAP is the faster nonlinear cousin, often preferred today for large sets, with the same caution about distances.
It is slow on large sets (hundreds of thousands of points need approximate versions) and has no notion of a new point: adding one means rerunning the whole map.
It is a tool for exploring and for spotting problems (a class of faults that never separates, mislabelled samples sitting in the wrong cluster), never evidence by itself. A claim found on a t-SNE map is checked in the original space, with clustering or a classifier.