mathematics//information theory//mutual information
Mutual information measures how many bits one variable tells about another, and it is what a sensor, a channel or a decoder is judged by when the question is how much a measurement could possibly reveal about what it measures. It is zero exactly when the two variables are independent, and it equals the whole entropy of one of them when the other determines it completely.
Mutual information measures how many bits one variable tells about another, and it is what a sensor, a channel or a decoder is judged by when the question is how much a measurement could possibly reveal about what it measures. It is zero exactly when the two variables are independent, and it equals the whole entropy of one of them when the other determines it completely.
It is defined from entropies, as the uncertainty about XXX minus the uncertainty that remains once YYY is known.
I(X;Y)=H(X)−H(X∣Y)I(X;Y) = H(X) - H(X \mid Y)I(X;Y)=H(X)−H(X∣Y)
The expression is symmetric, so YYY tells as much about XXX as XXX tells about YYY. A fair coin read through a channel that flips it one time in ten carries about 0.53 bits per reading instead of one, since 1−H(0.1)≈1−0.471 - H(0.1) \approx 1 - 0.471−H(0.1)≈1−0.47.
It counts every kind of dependence, linear or not, which is why it is preferred to correlation when the shape of the relation between a signal and what it encodes is unknown. The price is that it is hard to estimate from finite samples in many dimensions, and estimates from short recordings are biased upward.
Conditioning on side information changes the question. I(X;Y∣C)I(X;Y \mid C)I(X;Y∣C) is what YYY adds about XXX once CCC is already known, and it is the right quantity when a model brings knowledge of its own, a prior learned elsewhere, to the reading of a measurement.
Channel capacity is the mutual information between a channel's input and its output, maximised over the input distribution. How processing can only lose it is the data processing inequality.