ML//neural network//embedding//Word2Vec

Word2Vec is a method that learns one dense vector per word of a vocabulary from the company each word keeps in a large corpus, and it is the classic way to give a classical model (a classifier, a search engine, a clustering) word features that carry meaning. A shallow network is trained to predict a word from its neighbours, or the neighbours from the word, and the weights it ends with are the vectors: words that appear in the same contexts (*pump* and *compressor* in maintenance logs) end up close. GloVe reaches similar vectors another way, by factorizing the table of how often words co-occur; both came out of 2013 and 2014 and are the two names to know for **static word embeddings**.


Word2Vec is a method that learns one dense vector per word of a vocabulary from the company each word keeps in a large corpus, and it is the classic way to give a classical model (a classifier, a search engine, a clustering) word features that carry meaning. A shallow network is trained to predict a word from its neighbours, or the neighbours from the word, and the weights it ends with are the vectors: words that appear in the same contexts (pump and compressor in maintenance logs) end up close. GloVe reaches similar vectors another way, by factorizing the table of how often words co-occur; both came out of 2013 and 2014 and are the two names to know for static word embeddings.

Static is the word that separates them from what a transformer does. Word2Vec gives bank one vector, the same in the bank approved the loan and the bank of the river, an average of all its senses. A transformer starts from a table of vectors too (embedding matrix), but every layer of attention rewrites each vector with its context, so by the middle of the network the two banks have separate representations. That difference, one vector per word against one vector per word in its sentence, is the whole jump from the 2013 picture to the modern embedding.

A static embedding is a lookup table of meaning, frozen at training time.

It costs almost nothing to use (one row read per word) and still works as a feature for small models on a laptop or a PLC gateway, but it cannot tell senses apart and knows no word it never saw.

The famous arithmetic, king−man+woman≈queen\text{king}-\text{man}+\text{woman}\approx\text{queen}king−man+woman≈queen, shows that directions in the space carry relations. It holds for a few clean analogies and fails on many others, so it demonstrates structure in the space and proves no ability to reason.

Compared with a one-hot encoding, which gives every word its own axis and makes pump exactly as far from compressor as from banana, the dense vector is short (100 to 300 numbers) and puts related words near each other, measured with cosine similarity.

Words outside the training vocabulary get no vector at all, a real problem with part numbers, tag names and misspellings in plant data. Subword methods (fastText, and the tokenizers of modern models) fix this by building words from pieces.

To look at such a space, people project it to two dimensions with t-SNE; the picture comes after the vectors and does not produce them.