mathematics//decision theory//Markov decision process//POMDP
A POMDP (partially observable Markov decision process) is a Markov decision process in which the agent never sees the state, only noisy observations of it, so it must decide on a **belief**, the probability of each state given everything observed so far; it is the exact frame for a drone, a robot or a radar that acts on imperfect sensing. A chess player sees the whole board at every move (full observability); a drone sees a noisy part of its state through its sensors (partial observability), and most real systems are drones in this sense. This decision-theory sense is related to, and distinct from, the observability of control theory, which asks whether the state can be reconstructed from a history of outputs at all.
A POMDP (partially observable Markov decision process) is a Markov decision process in which the agent never sees the state, only noisy observations of it, so it must decide on a belief, the probability of each state given everything observed so far; it is the exact frame for a drone, a robot or a radar that acts on imperfect sensing. A chess player sees the whole board at every move (full observability); a drone sees a noisy part of its state through its sensors (partial observability), and most real systems are drones in this sense. This decision-theory sense is related to, and distinct from, the observability of control theory, which asks whether the state can be reconstructed from a history of outputs at all.
After doing aaa and observing yyy, the belief bbb is updated as
b′(s′)∝P(y∣s′)∑sP(s′∣s,a) b(s).b'(s') \propto P(y\mid s') \sum_{s} P(s'\mid s,a)\, b(s).b′(s′)∝P(y∣s′)s∑P(s′∣s,a)b(s).
The sum is the prediction (where the action takes you) and P(y∣s′)P(y\mid s')P(y∣s′) the correction by the measurement: this is the Bayes filter, of which the Kalman filter is the linear Gaussian case. A POMDP is equivalent to an MDP whose state is the belief, and that changes everything, because a belief is a whole distribution: continuous and huge. Solving a POMDP exactly is intractable in general, PSPACE-complete even with a finite horizon.
In exchange it sees something an MDP cannot: actions that do not approach the goal but make clear where you are, a detour to look around a corner or a stop to relocalize (active perception).
Without seeing the state you decide on beliefs, and acting as if the estimate were true works until the uncertainty matters.
Industry uses certainty equivalence with safety margins; a robot unsure whether a step is 30 cm or 3 cm away should not act as if it were at the mean.
Approximate solvers fill the middle ground: QMDP (which assumes the uncertainty vanishes after one step), point-based solvers that compute values only at sampled beliefs, and online tree search such as POMCP (Monte Carlo tree search over sampled states). They handle medium problems and remain niche bordering on research.
The belief is the right state, and a costly one. A particle filter can carry a belief with several hypotheses (a robot that may be in either of two identical corridors), and a planner that uses that whole cloud, instead of its mean, is what turns perception into decisions that seek information.