ML//RL//reward function
A reward function is the rule that gives a reinforcement-learning agent a number after each step or episode, telling it how well it did, and it is the only description of the goal the agent ever receives. The policy is trained to maximize the sum of these numbers over time, so writing the reward is writing the specification: whatever it leaves out, the agent ignores, and whatever it rewards by accident, the agent pursues.
A reward function is the rule that gives a reinforcement-learning agent a number after each step or episode, telling it how well it did, and it is the only description of the goal the agent ever receives. The policy is trained to maximize the sum of these numbers over time, so writing the reward is writing the specification: whatever it leaves out, the agent ignores, and whatever it rewards by accident, the agent pursues.
Where the answer can be checked, the reward is easy to write. A proof that verifies, a unit test that passes, a game won: one for success, zero otherwise.
R={1if the solution is correct0otherwiseR=\begin{cases}1 & \text{if the solution is correct}\\ 0 & \text{otherwise}\end{cases}R={10if the solution is correctotherwise
This binary reward is what made reasoning models trainable on mathematics and code (reasoning model). Most real goals are not like that: whether a maintenance decision was good, an explanation useful or a research direction promising has no immediate yes or no, and the reward has to be built.
Sparsity. A reward that arrives only at the end (a drone that gets a point only on landing at the target) gives almost no signal while the policy still crashes every time; learning stalls. Reward shaping adds intermediate rewards (getting closer, staying level) to guide it, and shaping done carelessly changes what is optimal; adding potential-based terms is the form that provably leaves the best policy unchanged.
Weighted criteria. A composite reward R=αRcorrect+βRsafe+γRefficientR=\alpha R_{\text{correct}}+\beta R_{\text{safe}}+\gamma R_{\text{efficient}}R=αRcorrect+βRsafe+γRefficient states several goals at once, and choosing α,β,γ\alpha,\beta,\gammaα,β,γ is choosing the trade-off. The weights do not solve the problem: each term can still be gamed, and a strong agent finds the corner where a cheap term outweighs the important one.
Misspecification. The written reward differs from the intended goal, and the agent optimizes the written one exactly (reward hacking, specification gaming). A learned reward, the reward model of language-model training, is a reward function fitted to human preferences, and it carries the same risk through its blind spots.
Credit. Even a correct reward at the end of a long episode leaves open which of the hundred actions earned it (credit assignment).
The reward is the specification, and convergence is only toward it.
An agent can train stably and converge to a policy that maximizes a wrong reward perfectly; the harder and more open the task, the more of the engineering goes into the reward and the checks around it, and the less into the learning algorithm.
In practice the reward is the first thing to review when a learned controller behaves strangely, and it is tested like code: by watching what the agent does when the reward is maximized in simulation before anything flies.