systems engineering//verification//fault injection
Fault injection is the deliberate introduction of failures into a system under test (a sensor that freezes, a delayed or lost message, a motor that delivers 70 % of its thrust, a supply that dips) to check that detection, reconfiguration and fail-safe behaviour work as specified; it is used in verification of autopilots, automotive controllers, industrial safety functions and distributed services, because the code that handles failures is the code that runs least and hides the most bugs. A failsafe that has never been exercised does not exist, whatever the design document says.
Fault injection is the deliberate introduction of failures into a system under test (a sensor that freezes, a delayed or lost message, a motor that delivers 70 % of its thrust, a supply that dips) to check that detection, reconfiguration and fail-safe behaviour work as specified; it is used in verification of autopilots, automotive controllers, industrial safety functions and distributed services, because the code that handles failures is the code that runs least and hides the most bugs. A failsafe that has never been exercised does not exist, whatever the design document says.
The faults can be injected at every rung of the test ladder. In software-in-the-loop the simulator corrupts what the code sees: a GPS that freezes at its last value, an IMU with a growing bias, 200 ms of extra latency, a radio link that drops 10 % of packets in bursts. In hardware-in-the-loop fault-insertion units short or open real sensor lines and disturb real buses. On a bench or a tethered vehicle a motor's command can be scaled down to see whether the control allocation compensates. In distributed software the same idea kills servers and partitions networks on purpose.
Faults are chosen from the analysis. Each failure mode found by hazard analysis and each fault the fault diagnosis design claims to detect becomes an injection case, with an expected response (detected within so many milliseconds, isolated to the right sensor, safe mode entered), so the test checks a written claim.
Detection is only half of the check. A residual that crosses its threshold is useless if the reconfiguration it triggers then fails: switching to a backup sensor with a different frame, a degraded mode never tuned, a return-home path that crosses the no-fly zone. Injection tests the whole chain to the safe state (fail-safe design).
Faults in a fleet need collective testing: fifty simulated drones losing the link at once show whether their individual fallbacks combine into a new hazard.
Injected faults must include the unpleasant ones: intermittent, slow drifts that stay under thresholds, two faults at once, and faults during transitions such as takeoff or a mode change, where real failures cluster.
Fault injection is one of the practices that make the pyramid of verification worth its cost; it is the test side of fault-tolerant control.