mathematics//information theory//mutual information//data processing inequality
The data processing inequality says that no computation on a measurement can carry more information about the source than the measurement itself, and it is the reason a better model can never stand in for a better sensor. If a source \(X\) produces a measurement \(Y\), and an output \(Z\) is computed from \(Y\) alone, then for any processing at all, deterministic or random, clever or not,
The data processing inequality says that no computation on a measurement can carry more information about the source than the measurement itself, and it is the reason a better model can never stand in for a better sensor. If a source XXX produces a measurement YYY, and an output ZZZ is computed from YYY alone, then for any processing at all, deterministic or random, clever or not,
I(X;Z)≤I(X;Y)I(X;Z) \le I(X;Y)I(X;Z)≤I(X;Y)
where III is the mutual information. The condition is that ZZZ sees XXX only through YYY (the three form a Markov chain).
What processing can do is make information usable. A decoder can pull out what was in the recording but buried in noise, and that gain can be enormous; what it cannot do is create information that never crossed the sensor.
A model can add knowledge, never evidence.
When the computation also uses side information CCC, such as a prior learned from other data, the inequality holds given that knowledge, I(X;Z∣C)≤I(X;Y∣C)I(X;Z \mid C) \le I(X;Y \mid C)I(X;Z∣C)≤I(X;Y∣C). The output may be far more plausible than the raw measurement, yet everything it says about this particular source beyond what CCC already said had to come through YYY.
It bounds every brain decoder. An output computed from a recording and a language model cannot say more about the person's state than the recording did, given the model, so a fluent reconstruction can report the prior as much as the brain.
Processing loses nothing exactly when ZZZ keeps everything in YYY that bears on XXX (a sufficient statistic). Compression, feature extraction and thresholding usually lose something, and the inequality says the loss can never be won back downstream.