ML//knowledge graph
A knowledge graph is a database that stores knowledge as entities (nodes) joined by typed relations (edges), each fact a triple such as *drug X, inhibits, protein Y*, and it is used where facts must be explicit, checkable and connected: answering questions that need several hops, enforcing what is known, and giving a language model something it can verify against. Search engines use one to answer *who directed this film*; a plant can keep one of its assets, where pump 3 *is part of* line 2, *is fed by* valve V-12 and *is maintained under* a given contract.
A knowledge graph is a database that stores knowledge as entities (nodes) joined by typed relations (edges), each fact a triple such as drug X, inhibits, protein Y, and it is used where facts must be explicit, checkable and connected: answering questions that need several hops, enforcing what is known, and giving a language model something it can verify against. Search engines use one to answer who directed this film; a plant can keep one of its assets, where pump 3 is part of line 2, is fed by valve V-12 and is maintained under a given contract.
Its strength is that every relation is stated and can be audited. The question which process could drug X affect? is answered by walking two edges (drug X inhibits protein Y, protein Y takes part in process Z), and the answer comes with the path that justifies it. Its weakness is the other side of the same coin: the graph knows only what someone entered or extracted, so it is incomplete, and keeping it current costs work. A clinical graph built on fixed relations will never propose the interaction nobody has recorded yet.
A language model and a graph fail in opposite ways. The model generalizes and proposes relations no one wrote down, and it can invent them (hallucination); the graph never invents and never proposes. Paired, the model suggests hypotheses and the graph confirms or rejects them against known relations and constraints, a neurosymbolic arrangement.
Feeding a graph to a model can be done in two ways. The plain one serializes the relevant triples as text into the context, which is what GraphRAG does with entities and community summaries extracted from documents. The deeper one learns vectors for nodes and relations (knowledge graph embeddings, or a GNN over the graph) and trains a projection into the model's input space; the difficulty is that the two representations were built differently and do not line up by themselves.
It is often confused with two neighbours. A vector database finds passages that are similar in meaning, with no notion of which entity relates to which and how; a world model predicts how a situation evolves under actions, where a graph only records how things stand. A graph can feed a world model, never replace it.