systems engineering//hazard analysis

Hazard analysis is the systematic search for what can go wrong in a system, how it could happen and how bad it would be, carried out before the requirements are written so that safety becomes part of what is specified; it is used in every safety-related project, from a chemical plant to a drone to a car's braking system, and its output is the list of hazards each safety function must address and how much each one must reduce the risk. Without it, requirements describe what the system should do on a good day, and the bad days are discovered in the field.


Hazard analysis is the systematic search for what can go wrong in a system, how it could happen and how bad it would be, carried out before the requirements are written so that safety becomes part of what is specified; it is used in every safety-related project, from a chemical plant to a drone to a car's braking system, and its output is the list of hazards each safety function must address and how much each one must reduce the risk. Without it, requirements describe what the system should do on a good day, and the bad days are discovered in the field.

Risk is ranked as probability times severity, and that product sets priorities: a frequent, minor event and a rare, catastrophic one may deserve similar attention, and the methods demand more evidence and more independence where the risk is highest. Several methods look at the system from different sides, and a serious project uses more than one.

HAZOP walks through a process diagram line by line with guide words (no, more, less, reverse, other than) applied to each parameter: what happens with no flow in this line, more pressure in this vessel? It is the standard method of the process industry, done by a team around a drawing.

FMEA starts from components: for each part, how can it fail, what is the effect, how would we detect it? It is bottom-up, exhaustive and good at single failures, and it misses combinations and failures without a broken part.

Fault tree analysis starts from one undesired event and works downward through AND and OR gates to the basic failures that cause it, which lets probabilities be combined and shows where a single point of failure sits.

STPA treats safety as a control problem: accidents come from unsafe control actions (a command given too late, at the wrong time, or not at all) in a loop of software, hardware and people. It finds hazards that arise with nothing broken, which is why it is preferred where software and operators interact.

The analysis is only as good as its imagination, so it is done by people who know the system, early, and again whenever the system changes. Each hazard found becomes a requirement with a required integrity level, which is what ties hazard analysis to the IEC 61508 lifecycle and to an automotive HARA under ISO 26262.

A hazard caused by no fault at all (a perception function misreading a scene) is the territory of SOTIF; a hazard caused by an attacker belongs to security analysis, which increasingly runs alongside. The analysis opens the functional safety process and sits at the start of systems engineering.