systems engineering//hazard analysis//FMEA

FMEA, failure mode and effects analysis, is a bottom-up method of hazard analysis that goes through a system component by component and asks, for each one, how it can fail, what each failure does to the system, and how it would be noticed; it is used in automotive and aerospace design, in process plants and in maintenance planning, and its output is a table that tells the team which failures need a design change, a diagnostic or a spare. In Spanish-speaking plants it is the AMFE.


FMEA, failure mode and effects analysis, is a bottom-up method of hazard analysis that goes through a system component by component and asks, for each one, how it can fail, what each failure does to the system, and how it would be noticed; it is used in automotive and aerospace design, in process plants and in maintenance planning, and its output is a table that tells the team which failures need a design change, a diagnostic or a spare. In Spanish-speaking plants it is the AMFE.

A row of the table is one failure mode of one part. For the bearing of a cooling pump: the mode is wear with rising vibration, the local effect is heat and noise, the system effect is a pump that seizes and a line that loses cooling, the cause is lubrication loss or misalignment, and the current detection is a monthly vibration round. Each row gets three scores from 1 to 10: severity of the effect, likelihood of occurrence, and how poorly the current controls would detect it before it bites. Traditionally they were multiplied into a risk priority number, and rows were worked from the top.

The product of the scores hides the catastrophes.

A severity of 10 with occurrence 2 and detection 5 gives the same RPN of 100 as a harmless but frequent and invisible failure, and the scales are ordinal, so multiplying them has no meaning. The 2019 AIAG-VDA handbook of the automotive industry replaced the RPN with an action-priority table that looks at the three ratings together, so that a high severity is never averaged away.

It is exhaustive about single failures and blind to the rest. Each row assumes one part fails and everything else works, so combinations (a sensor stuck while its backup is in maintenance), failures shared by redundant channels (common-mode failure) and hazards with nothing broken (software doing what it was told) escape it; those are the work of fault tree analysis, of STPA and of SOTIF.

Its detection column joins design to diagnosis. A failure mode with high severity and poor detection is a request for a sensor, a residual or a test (fault diagnosis); in maintenance the same table, with the time from detectable to failed added (P-F interval), decides the inspection interval.

It costs team time. A competent FMEA of a drone's power and propulsion runs to hundreds of rows written by people who know the parts, and a table filled in to satisfy an audit, never revisited after the design changed, is the common way it fails.