mathematics//decision theory//Markov decision process//value function
A value function is a function that gives, for each state of a sequential decision problem, the total discounted reward expected from that state onward under a given policy, and it is how a planner or a learning agent compares situations whose payoff lies in the future: a battery level, a wear level, a position on a map. Control engineers know it as the **cost-to-go**, the same quantity counted as cost instead of reward.
A value function is a function that gives, for each state of a sequential decision problem, the total discounted reward expected from that state onward under a given policy, and it is how a planner or a learning agent compares situations whose payoff lies in the future: a battery level, a wear level, a position on a map. Control engineers know it as the cost-to-go, the same quantity counted as cost instead of reward.
For a policy π\piπ in a Markov decision process, the state value is
Vπ(s)=E[∑t=0∞γtR(st,at) ∣ s0=s],V^{\pi}(s) = \mathbb{E}\Big[\sum_{t=0}^{\infty} \gamma^{t} R(s_t,a_t) \;\Big|\; s_0=s\Big],Vπ(s)=E[t=0∑∞γtR(st,at)s0=s],
the reward of every future step, discounted by γt\gamma^tγt and averaged over where the transitions may lead. The action value Qπ(s,a)Q^{\pi}(s,a)Qπ(s,a) is the same with the first action fixed to aaa, and the optimal values V∗V^*V∗ and Q∗Q^*Q∗ are those of the best policy, the ones the Bellman equation characterizes. A delivery drone at 40% battery two kilometres from base has a value that already contains the chance of finishing the round and the risk of a forced landing; a drone at the same battery beside the base is worth more, though neither has earned anything yet.
It is distinct from the reward and from the policy. The reward is what one step pays, the value is what the future is expected to pay from here, and the policy is what to do; with Q∗Q^*Q∗ in hand the optimal policy is just the best action in each state.
In practice it is the thing that is approximated, because representing it state by state explodes with the dimension (curse of dimensionality). Value iteration keeps it as a table; LQR has it in closed form as a quadratic x⊤Pxx^\top P xx⊤Px given by the Riccati equation; deep reinforcement learning fits a neural network to it, which in actor-critic methods is called the critic.
A value function learned from data is only as good as the states it was trained on. Q-learning estimates it from visited transitions, so states never visited keep their initial guess, and a policy that trusts those guesses can head confidently into the unknown.