ML//evaluation//BLEU

BLEU (bilingual evaluation understudy) is an automatic metric that scores a machine translation by how many of its short word sequences also appear in one or more human reference translations, and it is how translation systems have been compared on a test set since 2002, cheaply and reproducibly, without paying people to read every output. It counts the n-grams of the candidate (single words up to sequences of four) that occur in the references, clips repeated matches so that writing *the the the* earns nothing, takes the geometric mean of the four precisions, and multiplies by a **brevity penalty** that punishes outputs shorter than the references.


BLEU (bilingual evaluation understudy) is an automatic metric that scores a machine translation by how many of its short word sequences also appear in one or more human reference translations, and it is how translation systems have been compared on a test set since 2002, cheaply and reproducibly, without paying people to read every output. It counts the n-grams of the candidate (single words up to sequences of four) that occur in the references, clips repeated matches so that writing the the the earns nothing, takes the geometric mean of the four precisions, and multiplies by a brevity penalty that punishes outputs shorter than the references.

BLEU=BP⋅exp⁡ ⁣(14∑n=14log⁡pn),BP=min⁡ ⁣(1,  e 1−r/c)\text{BLEU}=\text{BP}\cdot\exp\!\left(\tfrac14\sum_{n=1}^{4}\log p_n\right),\qquad \text{BP}=\min\!\left(1,\;e^{\,1-r/c}\right)BLEU=BP⋅exp(41​n=1∑4​logpn​),BP=min(1,e1−r/c)

Here pnp_npn​ is the clipped share of the candidate's n-grams found in the references, ccc the candidate's length and rrr the reference length; the score runs from 0 to 1, usually reported times 100.

BLEU measures overlap of words and is blind to whether the meaning is right.

A translation that says the same thing with other words scores low, and one that copies the reference's phrasing while reversing a negation can score high. It is meaningful as an average over a whole test set and for comparing systems on the same references, and nearly meaningless for judging one sentence.

It is a precision metric (how much of what the system wrote is in the references); its sibling for summaries, ROUGE, leans on recall (how much of the reference the system covered). Both share the same blindness to paraphrase.

Numbers are only comparable with the same references, the same tokenization and the same implementation; papers once reported BLEU computed differently and got scores several points apart on the same output, which is why a standard implementation (sacreBLEU) is used today.

For instructed language models it is a weak measure: a helpful answer has no single reference to overlap with. Learned metrics, human evaluation or a judging model replace it for open generation, each with its own biases (benchmark).

It stays useful where outputs really are constrained, such as translating fixed-format maintenance messages or alarm texts between languages, where the wording is nearly determined and overlap is a fair proxy.