systems engineering//fail-safe design

Fail-safe design is the practice of deciding in advance what a system does when one of its parts fails or misses a deadline, and making that default behaviour lead to a safe state; it is used in every system where the normal path can break, from a planner that does not finish in time to a radio link that drops, and designing it matters as much as designing the algorithm it backs up. A mobile robot's planner normally replans at 10 Hz; when one cycle overruns, the controller keeps following the last valid plan, and when the plan runs out it stops or holds position. That rule, written before anything went wrong, is the fail-safe.


Fail-safe design is the practice of deciding in advance what a system does when one of its parts fails or misses a deadline, and making that default behaviour lead to a safe state; it is used in every system where the normal path can break, from a planner that does not finish in time to a radio link that drops, and designing it matters as much as designing the algorithm it backs up. A mobile robot's planner normally replans at 10 Hz; when one cycle overruns, the controller keeps following the last valid plan, and when the plan runs out it stops or holds position. That rule, written before anything went wrong, is the fail-safe.

The core choice is the safe state and how to reach it. For a factory robot it is usually stopped with brakes applied; for a valve it depends on the process (a fuel valve fails closed, a cooling-water valve fails open); for a drone it depends on altitude, battery and where it is: hold, return home or land now. Many mechanisms reach it without any software at all: a spring-return valve, a brake held off by power so that losing power applies it, a dead-man switch.

The fallback must not depend on what failed. A safe-stop command that travels through the same network, computer or software as the function it protects fails together with it. That is why immediate safety stays local (a robot's laser scanner stops it without asking a server 100 ms away), why a watchdog timer in independent hardware forces the safe state when the controller stops responding, and why a safety instrumented system is wired apart from the control system.

Missing a deadline is a failure mode of its own. A planner or an optimizer with no guaranteed run time (MPC computation, a search) needs a defined answer when it is late: the last valid plan, an anytime algorithm's best so far, or a conservative manoeuvre. Real-time computing bounds the cases where lateness can be ruled out.

In a fleet, the collective effect of the fallbacks matters. Fifty drones that all lose the link and return home along the same route create a new hazard; graceful degradation is designed for the group (what each does on link loss, what an isolated subgroup does, how a unit rejoins), not only for one vehicle.

Every fallback needs a test that forces it to run, in simulation and on the bench, which is the job of fault injection.

Fail-safe and fail-operational are different requirements. A train can stop; an aircraft in flight cannot, and neither can an electric barrier that contains only while it is energized, so they need redundancy that keeps functioning after a fault (redundancy, fault-tolerant control), which costs far more.

A learned or complex controller is often made certifiable by this pattern: a simple verified monitor watches it and switches to a verified fallback when the system leaves its safe envelope (Simplex architecture). Part of systems engineering, next to functional safety.