systems theory//complex system//cascading failure
A cascading failure is a failure that spreads through a network because the load or the work of each failed component shifts onto its neighbours, which overload and fail in turn, and it is how a local fault becomes a regional blackout or a stopped factory. In August 2003 in the northeast of the United States and Canada, a few transmission lines sagged into trees and an alarm system stopped warning the operators; within hours the outage had left around 50 million people without power. No individual failure explained the size of the result; the coupling did.
A cascading failure is a failure that spreads through a network because the load or the work of each failed component shifts onto its neighbours, which overload and fail in turn, and it is how a local fault becomes a regional blackout or a stopped factory. In August 2003 in the northeast of the United States and Canada, a few transmission lines sagged into trees and an alarm system stopped warning the operators; within hours the outage had left around 50 million people without power. No individual failure explained the size of the result; the coupling did.
The useful number is the branching ratio RRR, the mean number of new failures each failure causes. If RRR stays below one a cascade dies out, with an expected total size of
E[S]=11−R,R<1,\mathbb E[S]=\frac{1}{1-R},\qquad R<1,E[S]=1−R1,R<1,
so a network at R=0.5R=0.5R=0.5 loses two components per initial fault on average, one at R=0.9R=0.9R=0.9 loses ten, and as RRR approaches one the expected size blows up; at or above one the cascade may never stop (branching process). Load is what moves RRR: the same grid that shrugs off a line trip at night can cascade on a hot afternoon, when every line runs close to its limit and has little room to absorb a neighbour's flow.
A cascade is an unstable eigenvalue in network form.
Linearize how load is redistributed after a failure: if the resulting matrix has spectral radius above one, each failure triggers more than one on average, which is the stability question of a dynamical system asked of a network.
Factories cascade too. A line with small buffers between machines turns a two-minute stop at one station into an empty buffer downstream and a full one upstream, and the stop propagates along the whole line; buffers are firebreaks paid for in inventory.
The defences are margins and cuts. Grids are operated so that the loss of any single element leaves no other overloaded (the N-1 criterion), shed load deliberately before a cascade can, and split into islands when it starts; software systems use timeouts, circuit breakers and back-pressure for the same reason.
Cascade sizes are heavy-tailed: most events stay small and a few are enormous, so the historical average of blackouts underestimates the risk of the big one (heavy-tailed distribution). A cascade spreads one failure along the couplings, while a common-mode failure strikes many components at once from a shared cause, and the two call for different defences.