mathematics//information theory//KL divergence

The KL divergence (Kullback-Leibler divergence, or **relative entropy**) is a measure of how far one probability distribution is from a reference one, read as the cost of believing \(q\) when the truth is \(p\), and it is used to compare a model's predictions with data and the data a deployed model now sees with the data it was trained on. In bits, it is the extra code length per symbol paid by a code designed for \(q\) when the symbols actually follow \(p\).


The KL divergence (Kullback-Leibler divergence, or relative entropy) is a measure of how far one probability distribution is from a reference one, read as the cost of believing qqq when the truth is ppp, and it is used to compare a model's predictions with data and the data a deployed model now sees with the data it was trained on. In bits, it is the extra code length per symbol paid by a code designed for qqq when the symbols actually follow ppp.

DKL(p ∥ q)=∑xp(x)log⁡2p(x)q(x)D_{\mathrm{KL}}(p\,\|\,q)=\sum_x p(x)\log_2\frac{p(x)}{q(x)}DKL​(p∥q)=x∑​p(x)log2​q(x)p(x)​

Each term weighs the log ratio by how often the truth produces that value. A value the truth produces often and the belief thinks rare costs a lot; a value the belief expects and the truth never produces costs nothing in this direction. The sum is zero when the two distributions are equal and positive otherwise. A pump that runs half the time in production, p=(0.5,0.5)p=(0.5,0.5)p=(0.5,0.5), described by a model trained when it ran 90 % of the time, q=(0.9,0.1)q=(0.9,0.1)q=(0.9,0.1), gives D(p∥q)≈0.74D(p|q)\approx0.74D(p∥q)≈0.74 bits, and the reverse D(q∥p)≈0.53D(q|p)\approx0.53D(q∥p)≈0.53 bits.

It is asymmetric and it is not a distance.

The two directions D(p∥q)D(p|q)D(p∥q) and D(q∥p)D(q|p)D(q∥p) differ in general and the triangle inequality fails, so writing it down forces a choice of which side is the truth. A monitor that needs a symmetric number adds the two directions.

The cross-entropy that trains classifiers and language models differs from it by the entropy of the data: H(p,q)=H(p)+D(p∥q)H(p,q)=H(p)+D(p|q)H(p,q)=H(p)+D(p∥q). Mutual information is a KL divergence too, between the joint distribution of two variables and the product of their marginals, which is why it is zero exactly under independence.

Drift monitors use it on binned features. The vibration RMS of a pump is cut into ten bins, the fraction of recent data in each bin is pkp_kpk​ and of training data qkq_kqk​, and the population stability index ∑k(pk−qk)ln⁡(pk/qk)\sum_k(p_k-q_k)\ln(p_k/q_k)∑k​(pk​−qk​)ln(pk​/qk​) is exactly D(p∥q)+D(q∥p)D(p|q)+D(q|p)D(p∥q)+D(q∥p) in nats; by convention below 0.1 is stable and above 0.25 a serious shift (data drift).

An empty bin breaks it. Where q(x)=0q(x)=0q(x)=0 and p(x)>0p(x)>0p(x)>0 the divergence is infinite, so histograms are smoothed or sparse bins merged, and estimates from a few hundred samples come out biased upward: a small nonzero value on a short window is often noise.

In RLHF it is the leash: the reward the model maximizes is the reward model's score minus a multiple of the KL divergence between the trained policy and the starting one, so the model may improve its answers but not drift into text the reward model was never trained to judge. The penalty constrains the distribution and nothing more; it guarantees neither safety nor correctness.

Fitting a distribution by minimizing it behaves differently in each direction. Minimizing D(p∥q)D(p|q)D(p∥q) over qqq makes qqq spread to cover every region where the data live; minimizing D(q∥p)D(q|p)D(q∥p), as variational methods do, lets qqq lock onto one mode and ignore the others.