OS//concurrency//mutex//priority inversion
Priority inversion is the situation in a priority-scheduled system where a high-priority task waits for a resource held by a low-priority task, while a medium-priority task that needs no resource keeps the CPU and stops the low one from finishing, so the most urgent task is delayed by the least relevant one for an unbounded time. It is the classic way a real-time schedule that looked correct on paper misses its deadlines in the field.
Priority inversion is the situation in a priority-scheduled system where a high-priority task waits for a resource held by a low-priority task, while a medium-priority task that needs no resource keeps the CPU and stops the low one from finishing, so the most urgent task is delayed by the least relevant one for an unbounded time. It is the classic way a real-time schedule that looked correct on paper misses its deadlines in the field.
Three tasks make it happen. Low takes a mutex to update shared data. High wakes up, needs the same mutex, and blocks: that much is normal and short, since Low will release it in microseconds. Then Medium becomes ready; it outranks Low, so it runs, for as long as it likes. High, the most important task in the system, is now waiting on Medium, a task with which it shares nothing.
The canonical incident is the Mars Pathfinder reset of 1997. On the lander's VxWorks system a low-priority meteorological task held the mutex of the shared information bus, a medium-priority communications task ran long, and the high-priority bus management task missed its cycle; a watchdog timer saw the bus task had not run and reset the whole computer, several times. Engineers reproduced it on a replica on Earth from the system's event trace and fixed it by uploading a change that switched on priority inheritance for that mutex, an option the operating system already had but which was off by default.
Priority inheritance is the standard cure: while Low holds a resource that High is waiting for, Low temporarily runs at High's priority, so Medium cannot preempt it; when Low releases the mutex it drops back. The blocking time of High becomes bounded by one critical section of a lower task, a term that response-time analysis can add.
The priority ceiling protocol goes further: each mutex carries the priority of its most urgent user, and a task that takes it is raised to that level at once, which also prevents deadlocks between mutexes, at the price of raising priorities even when nobody is waiting.
It is a reason to choose a real-time operating system whose mutexes support inheritance, and to keep critical sections short. A cyclic executive cannot suffer it, having no priorities to invert.