mathematics//game theory//mixed strategy

A mixed strategy is a probability distribution over a player's actions, from which the action is drawn afresh each time the game is played, and it is used wherever being predictable can be exploited by someone watching: a patrol route, a spot-check schedule, the inspection of pipelines or containers. The opposite, always taking the same action in a given situation, is a **pure strategy**.


A mixed strategy is a probability distribution over a player's actions, from which the action is drawn afresh each time the game is played, and it is used wherever being predictable can be exploited by someone watching: a patrol route, a spot-check schedule, the inspection of pipelines or containers. The opposite, always taking the same action in a given situation, is a pure strategy.

The book's inspection game shows the gain in numbers. An inspector can watch one of two gates; behind A lies something worth 2, behind B something worth 1, and an intruder takes the value of whichever gate it enters if the inspector is elsewhere. Watching A always sends the intruder to B for a gain of 1; watching B always concedes 2. Watching A with probability ppp leaves the intruder 2(1−p)2(1-p)2(1−p) at A and ppp at B, and the inspector picks ppp so that both are equal.

2(1−p)=p⟹p=23,intruder’s gain=232(1-p) = p \quad\Longrightarrow\quad p=\tfrac{2}{3},\qquad \text{intruder's gain}=\tfrac{2}{3}2(1−p)=p⟹p=32​,intruder’s gain=32​

At that ppp the intruder is indifferent between the gates, and no choice it makes earns more than 2/32/32/3. The indifference is the computing rule in general: at the optimum, every action the opponent might use pays it the same, and the equations that say so give the probabilities.

Unpredictability pays when you are watched. The same single inspector cuts the intruder's gain from 1 to 2/32/32/3, a third less loss for no extra hardware, only by refusing to be predictable.

The randomness has to be real. A schedule from a fixed seed, a pattern tied to shift changes, or an operator who avoids gate B on cold nights leaks the draw, and the guarantee goes with it. The distribution itself can be published (the minimax theorem says so for zero-sum games); the draw comes from a source the opponent cannot model.

Against nature there is nothing to hide from. Wind, rain and wear do not study your schedule, so the best decision is the deterministic one that minimizes expected cost under the belief (decision theory), and randomizing only adds variance. Randomness in reinforcement learning has a third purpose: there it buys information about actions whose value is unknown (exploration-exploitation trade-off), and here it denies information to an opponent.

Mixed strategies are the leader's tool in a Stackelberg game, and many Nash equilibria of general games exist only in mixed form (matching pennies has no pure one).