ML//neural network//embedding//platonic representation hypothesis
The platonic representation hypothesis is the claim that neural networks trained on different data, objectives and senses drift, as they grow, toward the same way of judging which things are similar, and it is used as a reason to expect what one model learned to carry over to another model, another modality or another kind of system. Huh, Cheung, Wang and Isola stated it in 2024 and measured it by comparing which inputs each model places near each other: two models agree when the nearest neighbours of an input in one embedding space are also its nearest neighbours in the other, and that agreement rose with model size, within vision and between vision and language models.
The platonic representation hypothesis is the claim that neural networks trained on different data, objectives and senses drift, as they grow, toward the same way of judging which things are similar, and it is used as a reason to expect what one model learned to carry over to another model, another modality or another kind of system. Huh, Cheung, Wang and Isola stated it in 2024 and measured it by comparing which inputs each model places near each other: two models agree when the nearest neighbours of an input in one embedding space are also its nearest neighbours in the other, and that agreement rose with model size, within vision and between vision and language models.
The argument behind it is that every model is fitting the same world. If the data are views of one underlying reality, models that predict them well have to capture its statistics, and their geometries converge on it. The formal version holds for exact, invertible views of one world and is weaker for views that are noisy, partial or carry different information, which is the case of real senses.
It is a hypothesis with a measured trend. The convergence is partial, the alignment metrics are relative, and two models that differ in what they can see (a caption cannot hold everything in an image) cannot converge fully.
Its practical use is transfer. If two spaces share structure, a small map fitted on a few paired examples can carry one into the other, which is what transfer learning and the alignment of brain recordings to language model features rely on.
Sharing neighbours is weaker than sharing a space. Two representations related by a map that stretches some directions and shrinks others can agree on what is near what while disagreeing on every distance.