control//resilient control

Resilient control is the design of control systems that keep an acceptable level of operation in the face of faults, attacks and uncertainty together, and recover afterwards; it is what infrastructure operators mean when they ask how a grid, a water network or a drone fleet behaves on its worst day. No design prevents everything, so resilience is about what happens after the blow: absorb it, adapt, recover.


Resilient control is the design of control systems that keep an acceptable level of operation in the face of faults, attacks and uncertainty together, and recover afterwards; it is what infrastructure operators mean when they ask how a grid, a water network or a drone fleet behaves on its worst day. No design prevents everything, so resilience is about what happens after the blow: absorb it, adapt, recover.

It is drawn as a resilience curve, performance against time: it dips with the incident, holds on a degraded plateau and climbs back. The area lost under the curve is the damage. A resilient system does not promise never to fall. It promises a floor (the drone lands instead of crashing, the plant produces at 60 % instead of stopping) and a bounded recovery. That is the difference from fault tolerance, which aims to keep the function through anticipated accidental faults, and from robustness, which keeps stability for a described family of plants; resilience joins both with security, because a fault with intent is designed to stay inside the detector's band (stealthy attack).

A system that has never practised its degraded mode does not have one.

Graceful degradation is designed and tested behaviour: what each drone does on link loss (hold, return, land), what an isolated subgroup does, how a node rejoins. Fifty drones returning home at once along the same route are a new problem the individual failsafe never saw.

Resilient state estimation keeps the estimate sound when a few sensors are attacked. Minimizing the sum of absolute errors instead of squares, x^=arg⁡min⁡x∑i∣yi−Cix∣\hat x=\arg\min_x\sum_i|y_i-C_ix|x^=argminx​∑i​∣yi​−Ci​x∣, lets a handful of huge errors weigh little, so with enough redundancy the estimator stops listening to the attacked sensors instead of averaging them in (robust statistics).

Defences rest on what the attacker does not know or control: independent physical measurements outside the network, many cross-relations (forging one value is easy, forging twenty consistent with physics is not), segmented networks, and a hard-wired safety instrumented system that trips on physical limits without asking the control software.

In fleets the averaging that builds consensus is hijacked by one bad node; resilient algorithms discard the most extreme neighbour values before averaging (resilient consensus).

Maturity: designed and tested degraded modes are good industry practice; resilient control as a formal theory is research (cyber-physical security).