industrial//reliability//availability

Availability is the fraction of time a repairable system is able to perform its function, and it is the figure that production planning, service contracts and grid operators actually care about: a machine that fails often but is repaired in minutes can be more available than one that fails rarely and waits weeks for a part. It combines how often units fail with how long they stay down, which reliability alone does not say. For a simple repair cycle it is the mean time up divided by the mean time of a whole cycle.


Availability is the fraction of time a repairable system is able to perform its function, and it is the figure that production planning, service contracts and grid operators actually care about: a machine that fails often but is repaired in minutes can be more available than one that fails rarely and waits weeks for a part. It combines how often units fail with how long they stay down, which reliability alone does not say. For a simple repair cycle it is the mean time up divided by the mean time of a whole cycle.

When a machine passes through intermediate states, a small Markov chain computes it. Let a machine be healthy (H), degraded (D) or failed (F), and each day move between them with probabilities that depend only on today's state: a healthy machine degrades with 2 % probability, a degraded one fails with 10 %, a failed one is repaired with 50 %. As a transition matrix, rows summing to one,

P=[0.980.02000.900.100.500.5],π⋆=π⋆P.P=\begin{bmatrix}0.98&0.02&0\\ 0&0.90&0.10\\ 0.5&0&0.5\end{bmatrix},\qquad \pi^{\star}=\pi^{\star}P .P=​0.9800.5​0.020.900​00.100.5​​,π⋆=π⋆P.

The long-run share of days in each state is the row vector π⋆\pi^{\star}π⋆ that the matrix leaves unchanged, its stationary distribution. Here it comes out at about 80.6 % healthy, 16.1 % degraded and 3.2 % failed: an availability of 96.8 %.

Availability is a design variable as much as a measurement.

Each number in the matrix is a lever: faster detection of the degraded state, a spare kept on site that raises the repair probability, a condition-based intervention that sends degraded machines back to healthy before they fail. Rerunning the chain with the new probabilities prices each lever in hours of production before any money is spent.

How fast the chain forgets its starting state is set by its second eigenvalue, about 0.875 here: after a repair campaign or a bad batch of parts, the gap to the long-run mix shrinks by that factor each day, halving about every five days.

Memorylessness is a modelling choice. If the chance of failing depends on how long the machine has been degraded, the state is enlarged (degraded for one day, two days, more) until the dependence is inside it.

Redundancy raises availability only as far as the copies fail independently; a shared cause takes them down together and the parallel arithmetic overpromises (common-mode failure).