ML//RL//exploration-exploitation trade-off
The exploration-exploitation trade-off is the dilemma, faced by any system that learns from the outcomes of its own decisions, between acting on what it currently believes is best (**exploitation**) and trying other actions to learn what they are worth (**exploration**). It decides how a learning controller, a dispatcher or a recommender gathers its own data, and it is shared by reinforcement learning and adaptive control.
The exploration-exploitation trade-off is the dilemma, faced by any system that learns from the outcomes of its own decisions, between acting on what it currently believes is best (exploitation) and trying other actions to learn what they are worth (exploration). It decides how a learning controller, a dispatcher or a recommender gathers its own data, and it is shared by reinforcement learning and adaptive control.
Its root is that decision errors do not average out the way noise does, because a system only observes the result of what it did. A fleet dispatcher that sends robots down aisle A because aisle A looked faster in the first week never collects a single trip time for aisle B, and so never learns that B became the better route once the shelves were moved. The data are chosen by the policy that is being judged by them, which is why the book calls it decision feedback, or selection bias in decisions. A maintenance policy that always replaces a part at 2,000 hours has the same blindness: it never sees how long the part would have lasted.
A system that acts pays to learn.
Exploring costs reward now, and on hardware it can cost the hardware; exploiting costs knowledge, and a system that only exploits can stay locked on a worse option for good.
In RL the exploration is written into the algorithm. Q-learning only converges if every state-action pair keeps being visited, so an ε\varepsilonε-greedy agent acts at random a fraction ε\varepsilonε of the time. On a real plant exploring means breaking things, which pushes the learning into simulation.
In control the same dilemma is dual control, formalized by Feldbaum around 1960: a good regulator keeps the system still, and a still system teaches nothing about its parameters (persistent excitation). Practice injects small pseudo-random excitations during commissioning or maintenance windows, uses the natural manoeuvres (take-off, turns) and freezes the adaptation when no information arrives.
The principled answer is to explore when it pays. Value of information prices an observation by how much it can change later decisions, and a POMDP chooses information-gathering actions for that reason; a fixed exploration rate is the cheap approximation of the same calculation, and the multi-armed bandit is the stripped-down setting where better approximations (an optimism bonus, sampling from the posterior) are worked out.