ML//RL//temporal credit assignment

Credit assignment is the problem of deciding which of the many actions or components that led to an outcome deserve the credit or the blame for it, and every learning system that receives delayed or aggregate feedback has to solve it. A drone that crashes ninety seconds into a flight, a chess game lost at move sixty, a language model answer scored as a whole: the outcome is one number, the decisions that produced it are many.


Credit assignment is the problem of deciding which of the many actions or components that led to an outcome deserve the credit or the blame for it, and every learning system that receives delayed or aggregate feedback has to solve it. A drone that crashes ninety seconds into a flight, a chess game lost at move sixty, a language model answer scored as a whole: the outcome is one number, the decisions that produced it are many.

There are two versions. Temporal credit assignment spreads a delayed reward back over the sequence of actions that preceded it; structural credit assignment spreads an error over the parts of a system that computed it, which in a neural network is exactly what backpropagation does, giving each weight its share of the gradient.

Reinforcement learning attacks the temporal version with value estimates. A value function predicts how good each state is, and the temporal-difference error credits each step with the change in that prediction it caused, so a reward at the end flows backward one step at a time instead of being smeared evenly over the episode. The discount factor decides how far back credit reaches.

Without that machinery every action of an episode gets the same credit. Sequence-level rewards in language-model training have this defect: in PPO with a whole-response score every token receives the same signal, which is one argument for per-step scoring with process reward models and for methods that derive a per-token gradient, such as DPO.

Sparse rewards make it worse (reward function): the rarer and later the signal, the more actions share it and the more samples are needed before the right ones stand out.

Feedback is cheap; knowing what it was about is expensive.

Most of the sample inefficiency of reinforcement learning is credit assignment, and most of the tricks that make it practical (value functions, shaping, step-level scores, shorter episodes) are ways of making each piece of feedback point at fewer decisions.

The same problem appears outside learning. A production line misses its quota and five machines, two shifts and a supplier each contributed; a fleet mission fails and the cause is buried in a chain of handoffs. Root-cause analysis is credit assignment done by people, and it fails for the same reason: the outcome is observed once, and the counterfactual for each contributor is not.