systems engineering//single point of failure
A single point of failure is a component whose loss alone stops the whole system, and finding them is one of the first jobs of an architecture review; the term is used for hardware (one power supply, one pump), software (one server, one database), communication (one radio link) and people (the one engineer who understands the fleet manager). A warehouse fleet coordinated by one central server stops when the server or its network does; a fleet of drones that all localize against one shared map server goes blind with it.
A single point of failure is a component whose loss alone stops the whole system, and finding them is one of the first jobs of an architecture review; the term is used for hardware (one power supply, one pump), software (one server, one database), communication (one radio link) and people (the one engineer who understands the fleet manager). A warehouse fleet coordinated by one central server stops when the server or its network does; a fleet of drones that all localize against one shared map server goes blind with it.
The remedies cost something each time. Redundancy duplicates the component so another takes over (redundancy), at the price of hardware, of the switch-over logic and of the testing that logic needs. Distribution removes the centre altogether, at the price of optimality, debuggability and verification: agents deciding on partial, stale information produce failures that appear at four in the morning in the customer's warehouse. Degradation accepts the loss and defines what the system still does without the part (fail-safe design).
Removing a single point of failure is a trade, judged against what the failure costs. A centralized fleet keeps its single point on purpose because a centre can plan optimally, keeps one log and one truth, and is far easier to certify; the book's advice is to centralize while communications and computation allow and distribute only what must survive without the link or react faster than the round trip (fleet architecture).
Duplicates that share a cause are one point, however many copies exist. Three identical sensors frozen by the same ice or three copies of the same software with the same bug fail together (common-mode failure), which is why critical systems mix suppliers or designs.
Network structure hides them. A graph with a few highly connected hubs survives random failures well and collapses when a hub is hit (scale-free network); a shared map server is such a hub whether or not anyone drew it as one.
Fault tree analysis makes them visible: any basic event that reaches the top event through OR gates alone is a single point of failure.
The concept belongs to systems engineering and to the reliability side of distributed design (distributed systems).