ML//model//Large Concept Model

A Large Concept Model (LCM) is a language model that predicts the next sentence-level idea instead of the next token, working in a space where each sentence is one fixed-size vector, and it is a research line from Meta (December 2024) that asks whether a model reasons better when its unit of thought is closer to a human idea than to a word piece. Its text is first cut into sentences, each sentence is encoded by a frozen encoder (SONAR, which covers about 200 languages in text and speech) into one vector, the model predicts the vector of the next sentence, and a frozen decoder turns that vector back into words.


A Large Concept Model (LCM) is a language model that predicts the next sentence-level idea instead of the next token, working in a space where each sentence is one fixed-size vector, and it is a research line from Meta (December 2024) that asks whether a model reasons better when its unit of thought is closer to a human idea than to a word piece. Its text is first cut into sentences, each sentence is encoded by a frozen encoder (SONAR, which covers about 200 languages in text and speech) into one vector, the model predicts the vector of the next sentence, and a frozen decoder turns that vector back into words.

The interest is in where the reasoning happens. An ordinary LLM plans a paragraph one token at a time; an LCM plans it one sentence at a time, in a representation that is the same whatever the language, so in principle a model trained mostly in English writes in Spanish through the same concepts. The authors tried several ways to predict the next vector (plain regression, diffusion-style generation, a quantized space) at models of about 1.6 billion parameters, with promising results in summarization and cross-language generalization and no claim yet of beating comparable LLMs across the board.

Bytes, tokens and concepts are three sizes of the unit a model predicts, and none of them is a higher rung of intelligence.

A byte-level model (BLT) goes below the token to avoid the tokenizer; an LCM goes above it to plan in ideas. Each moves a cost and a weakness to another place.

The concept is only as good as the encoder that defines it. SONAR's vectors were trained for translation and similarity, so what counts as one idea, and what detail is lost when a sentence becomes 1,024 numbers, is decided by that frozen model.

Long, information-dense sentences (a specification line with six numbers in it) are where a single vector per sentence is most likely to blur, while a token model keeps every digit.

It belongs to the same family of attempts to reason in a latent space rather than in visible text, alongside work on models that think in hidden vectors instead of written chains of thought.

For an engineer it is a direction to watch: no production system ran on it when it was published, and the practical models remain token-based.