ML//GNN

A GNN (graph neural network) is a neural network that operates on data structured as a graph, updating a vector at each node by mixing it with the vectors of its neighbours at every layer, and it is used when the connections are part of the data: drones that talk to their radio neighbours, substations and lines of a power grid, pumps and pipes of a water network, machines on a line that pass parts to each other. A table, an image or a sequence would throw that structure away.


A GNN (graph neural network) is a neural network that operates on data structured as a graph, updating a vector at each node by mixing it with the vectors of its neighbours at every layer, and it is used when the connections are part of the data: drones that talk to their radio neighbours, substations and lines of a power grid, pumps and pipes of a water network, machines on a line that pass parts to each other. A table, an image or a sequence would throw that structure away.

hv(ℓ+1)=ϕ ⁣(W1hv(ℓ)+W2∑u∈N(v)hu(ℓ))h_v^{(\ell+1)}=\phi\!\left(W_1h_v^{(\ell)}+W_2\sum_{u\in N(v)}h_u^{(\ell)}\right)hv(ℓ+1)​=ϕ​W1​hv(ℓ)​+W2​u∈N(v)∑​hu(ℓ)​​

Here hv(ℓ)h_v^{(\ell)}hv(ℓ)​ is the vector of node vvv at layer ℓ\ellℓ, N(v)N(v)N(v) its neighbours and W1W_1W1​, W2W_2W2​ learned matrices shared by all nodes. This is message passing: each node asks its neighbours, mixes what it hears with what it knows, and repeats, so after LLL layers it knows its neighbourhood LLL hops away. Because the weights are common to all nodes, the same network serves a fleet of 10 drones or of 200 (weight sharing).

The most cited variant, the GCN (graph convolutional network), multiplies at each layer by the normalized adjacency D~−1/2(A+I)D~−1/2\tilde D^{-1/2}(A+I)\tilde D^{-1/2}D~−1/2(A+I)D~−1/2, with AAA the adjacency matrix and D~\tilde DD~ the degree matrix of A+IA+IA+I. That is one step of diffusion on the graph with learned weights in between, a direct relative of the graph Laplacian, whose eigenvectors are the graph's frequencies; a GNN filters signals in that basis, as an ordinary convolution filters a signal in the Fourier basis.

Depth is limited by over-smoothing: with many layers every node ends up with the same vector, as in diffusion and consensus. A few layers are typical.

Uses in production are niche but real: detecting faults and predicting flows in power and water networks, estimating travel times on road networks (deployed in map services), coordinating agents in a fleet. A water network with 300 pressure sensors can use its topology to localize a leak from the pressure drops at neighbouring nodes.

If the graph is small and fixed, adding the neighbours' averages as features to a classical model captures much of the effect; the GNN pays when the graph is large or changes. Its inductive bias is locality, and it is wrong when the influence that matters travels far.

Coupling a GNN to a language model needs a bridge, because node vectors live in a different space from the LLM's token embeddings. They pass through a projector trained on pairs of graphs and descriptions (or become special tokens), the same recipe that lets a vision encoder talk to an LLM; the simpler alternative is letting the LLM query the graph as a tool and read the answer as text. Labelled graph data is scarce, so weak supervision (labels produced by rules or heuristics, noisy but plentiful) is common in graph tasks such as flagging toxic molecules.

The acronym is unrelated to global nearest neighbour, an assignment rule of data association in tracking.