control//fault-tolerant control

Fault-tolerant control is the design of control systems that keep performing their function, perhaps degraded, when a sensor, an actuator, a processor or a piece of the plant fails, and it is what lets an airliner keep flying with a dead computer and a hexacopter keep flying with a dead motor. Every member of this family answers the same question (what happens to the loop the instant something breaks) and pays for its answer in hardware, weight, software complexity and testing.


Fault-tolerant control is the design of control systems that keep performing their function, perhaps degraded, when a sensor, an actuator, a processor or a piece of the plant fails, and it is what lets an airliner keep flying with a dead computer and a hexacopter keep flying with a dead motor. Every member of this family answers the same question (what happens to the loop the instant something breaks) and pays for its answer in hardware, weight, software complexity and testing.

There are two routes. The passive one designs a controller that withstands a set of faults without noticing them: a robust design whose margins cover a motor at 70 %, the way the planar drone keeps flying (badly) when its right motor weakens. It needs no diagnosis and reacts instantly, but it covers only the faults inside its uncertainty set and pays with conservatism all the time. The active route detects the fault, isolates it (fault diagnosis) and then reconfigures: changes sensor, actuator allocation, controller or mission (reconfiguration). It covers larger faults with less conservatism, but it is only as good as its diagnosis, and its switching logic is code that rarely runs.

A tolerated fault is a fault someone imagined in advance.

A passive design tolerates the faults inside the uncertainty set it was designed for; an active one tolerates the faults its diagnosis can isolate and its reconfiguration logic covers. Everything else falls through to the safe mode, so the list of faults considered, written early in a hazard analysis, is the real specification of the system.

The raw material is redundancy: physical (identical units compared), dissimilar (different designs against common causes) and analytical (a model acting as a second sensor, analytical redundancy). With three channels the median ignores a wild one (majority voting).

Fly-by-wire is the reference application: electronic flight controls where the pilot's stick commands computers, not cables, and the computers are redundant and dissimilar (different processors and software teams) so that one design error cannot take them all down.

A single safe mode designed with care (stop, land, close the valve) covers most faults in most products; reconfiguring in flight pays when stopping is dangerous or very expensive (fail-safe design).

Fault tolerance is about accidental failures. A failure with intent (a spoofed sensor, injected data) needs the same redundancy plus the assumption that the adversary knows your detector (resilient control, cyber-physical security).

Maturity: voting, redundant computers and degraded modes are industry in aviation, cars and process plants; online reconfiguration of the control law is niche outside aerospace.