ML//neural network//embedding//one-hot encoding

A one-hot encoding is a representation of a categorical value as a vector with one position per possible category, all zeros except a single 1 at the position of the value it holds, and it is the standard way to feed a category (a machine type, a shift, a fault code, a word) to a model that only computes with numbers. With three pump models A, B and C, model B becomes \((0,1,0)\); with a vocabulary of 50,000 words, a word becomes a vector of 50,000 numbers with one 1 in it.


A one-hot encoding is a representation of a categorical value as a vector with one position per possible category, all zeros except a single 1 at the position of the value it holds, and it is the standard way to feed a category (a machine type, a shift, a fault code, a word) to a model that only computes with numbers. With three pump models A, B and C, model B becomes (0,1,0)(0,1,0)(0,1,0); with a vocabulary of 50,000 words, a word becomes a vector of 50,000 numbers with one 1 in it.

Its virtue is that it assumes nothing. Writing the pump models as 1, 2 and 3 would tell a linear regression that C is three times A and that B sits halfway between them, an order the data never had; the one-hot vector keeps every category equally far from every other. Its defect is the same fact seen from the other side: pump is as far from compressor as from banana, so the representation carries no notion of similarity, and its length grows with the number of categories.

One-hot is sparse and says only which; an embedding is dense and says what it is like.

A transformer starts exactly there: the token ID is in effect a one-hot vector, and multiplying it by the embedding matrix just picks one row, the dense vector the network learned for that token.

A sparse vector has most of its entries at zero (one-hot, word counts, TF-IDF); a dense vector has most of them nonzero and spreads information across all coordinates (embedding, Word2Vec). Sparse vectors are stored as the list of their nonzero positions, which is why a 50,000-wide vector costs a handful of numbers.

A model's output layer often speaks one-hot in reverse: the training target for the next token is a one-hot vector, and the cross-entropy loss compares it with the probabilities the softmax produced.

With many categories (thousands of part numbers, sensor tags), one-hot columns explode the width of a table and most of them stay empty in any small dataset. The usual remedies are grouping rare categories, target encoding, or learning a small embedding per category.

In statistics the same encoding is called dummy variables, with one column dropped so that the columns are not perfectly redundant with the intercept.