Credit assignment · Grey Matter

Credit assignment is the problem every learning system faces of working out which of its many connections, among all those active before an outcome, deserve to change and by how much, and how the brain solves it is one of the open questions that separate brains from today's artificial networks.


Credit assignment. Credit assignment is the problem every learning system faces of working out which of its many connections, among all those active before an outcome, deserve to change and by how much, and how the brain solves it is one of the open questions that separate brains from today's artificial networks.

Artificial neural networks solve it with backpropagation. After each example, the error at the output is sent backwards through the network, layer by layer, using the same weights as the forward pass and the derivative of each unit's activation, so every weight receives exactly its share of the blame. It works extremely well, and almost every step of it is hard to map onto neurons. Synapses transmit in one direction, so a backward pass would need a separate pathway whose weights mirror the forward ones (the weight transport problem); neurons send discrete spikes, while the algorithm needs graded error values and derivatives; and the brain learns while it acts, continuously, with no global controller alternating a forward and a backward phase.

The mainstream view is that the brain does not run backpropagation as written. A review in 2020 argued it may approximate its effect, for example by using feedback connections to carry differences in activity that act as local error signals.

The weight symmetry turned out to be less necessary than thought. In feedback alignment, error sent back through fixed random weights still trains a network, because the forward weights come to align with the feedback.

Neuromodulators supply a third factor. Dopamine neurons fire to the difference between expected and received reward, and learning rules in which a synapse marks itself when its two cells are active together (an eligibility trace) and changes only if a later dopamine or other modulator signal arrives can assign credit across seconds.

Attention and rhythms are proposed helpers. Prefrontal cortex, basal ganglia and thalamic circuits that select what is processed, and oscillations that time when plasticity can happen, are candidate ways of restricting which synapses learn, with less direct evidence.

The brain has to learn locally what backpropagation computes globally.

Each synapse sees only its own two cells and whatever modulators wash over it, and the open question is how that local view adds up to learning as good as it is.

Questions: How could dopamine tell a synapse, seconds later, that it helped? Dopamine neurons fire more when an outcome is better than expected and less when it is worse, a reward prediction error broadcast to large parts of the brain. In three-factor learning rules, a synapse whose two neurons were active together sets a temporary chemical tag, an eligibility trace that fades over about a second or more, and the synapse changes only if dopamine or another modulator arrives while the tag is still there. The tag says which synapses were involved and the modulator says whether it went well, which assigns credit across the delay between an action and its result, with experimental support in the striatum, hippocampus and cortex. Why does the brain probably not learn by backpropagation as artificial networks do? Backpropagation sends an exact error backwards through the same connections used in the forward pass, which neurons cannot do directly: synapses transmit one way, so the backward path would need separate connections mirroring every forward weight. It also needs graded error values and the derivative of each unit's response, separate phases for computing and for learning, and a controller alternating them, while neurons send discrete spikes and learn continuously as they act. Most researchers conclude that the brain uses local rules that approximate some of its effect, helped by feedback connections and neuromodulators, and the finding that error sent through random feedback weights still trains networks shows the symmetry requirement can be relaxed.