ML//evaluation//ROUGE
ROUGE is a family of automatic metrics that score a generated summary by how much of the word content of human reference summaries it recovers, and it is the standard cheap check for summarization systems, from news digests to the shift-handover summaries a model might write from operator logs. **ROUGE-N** counts the n-grams of the reference that appear in the candidate (ROUGE-1 for single words, ROUGE-2 for pairs); **ROUGE-L** looks for the longest common subsequence, words appearing in the same order without having to be adjacent, which rewards keeping the reference's structure.
ROUGE is a family of automatic metrics that score a generated summary by how much of the word content of human reference summaries it recovers, and it is the standard cheap check for summarization systems, from news digests to the shift-handover summaries a model might write from operator logs. ROUGE-N counts the n-grams of the reference that appear in the candidate (ROUGE-1 for single words, ROUGE-2 for pairs); ROUGE-L looks for the longest common subsequence, words appearing in the same order without having to be adjacent, which rewards keeping the reference's structure.
Where BLEU asks how much of the candidate is in the references, ROUGE leans the other way and asks how much of the reference the candidate covered. That fits summarization, where missing a key fact is the worst failure, and it is usually reported as recall, precision and their F1 together, so a summary cannot win by being long.
ROUGE rewards reusing the reference's words, and a wrong fact in the same words costs almost nothing.
A summary that says the pump was not restarted when the reference says the pump was restarted shares almost every word and scores well, and a faithful summary in fresh words scores badly. It ranks systems on average and flags gross failures; it cannot certify a single summary.
Factual errors are the failure that matters for summaries, and overlap metrics do not see them. Checks that compare the summary's claims against the source (by people, by a judging model, or by extracting entities and verifying each appears in the source) catch what ROUGE misses.
The score depends on preprocessing (stemming, stopwords, tokenization) and on how many references exist; numbers from different setups are not comparable.
Like BLEU, it was built for systems that copy or compress source sentences. Instructed models rephrase freely, so their ROUGE against a single reference understates their quality and overstates the differences between two good systems.
In a project it earns its place as a regression check: the same test set, the same references, run on every model version, where a sudden drop means something broke even if the absolute number means little (benchmark).