computer science//distributed systems//network partition
A network partition is a failure in which the network splits a distributed system into groups that can talk inside themselves but not to each other while every node keeps running, and it is the failure that forces the hardest choices in replicated databases and robot fleets alike. Each group sees the other go silent and, with no way to tell a crash from a cut link, keeps working convinced that the other has vanished.
A network partition is a failure in which the network splits a distributed system into groups that can talk inside themselves but not to each other while every node keeps running, and it is the failure that forces the hardest choices in replicated databases and robot fleets alike. Each group sees the other go silent and, with no way to tell a crash from a cut link, keeps working convinced that the other has vanished.
Partitions are ordinary events. A switch fails between two server racks; a drone flies behind a hangar and loses the radio for two seconds; a fleet spread over a mine splits into two groups out of each other's range (communication graph). In the graph view a partition is a disconnected communication graph, and its signature is exact: the second eigenvalue of the Laplacian, λ2\lambda_2λ2, drops to zero (spectral gap).
During a partition each group decides alone, and the decisions have to be reconciled when it heals.
That is the moment naive designs break: two groups that each reassigned the same task, two servers that each accepted a write (the split brain of database clusters), two halves of a fleet that each settled on their own meeting point.
In replicated data it is the P of the CAP theorem. Each side either keeps answering with what it knows, and the conflicts are merged later, or the minority side refuses until the majority is reachable again, as Raft does.
In a consensus protocol each connected piece converges to the average of its own members, so a split fleet never agrees on one value; the disagreement stops falling at the moment the graph splits.
It has to be designed for explicitly: what each agent does when it loses contact (hold, return, land), what an isolated subgroup may decide on its own, and how a node that rejoins catches up (resilient control).
It is one of three errors a fleet adds to those of a single robot, beside stale information and an inconsistent view (communication constraints), and all three travel up into assignment, estimation and control.