mathematics//similarity measure//Jaccard index
The Jaccard index is a similarity measure between two sets, the number of elements they share divided by the number of elements in either, and it is used wherever things are present or absent rather than measured: comparing the keywords of two documents, the alarms of two incidents, the parts of two bills of materials, the pixels of a predicted and a true defect mask. It runs from 0 (nothing in common) to 1 (the same set).
The Jaccard index is a similarity measure between two sets, the number of elements they share divided by the number of elements in either, and it is used wherever things are present or absent rather than measured: comparing the keywords of two documents, the alarms of two incidents, the parts of two bills of materials, the pixels of a predicted and a true defect mask. It runs from 0 (nothing in common) to 1 (the same set).
J(A,B)=∣A∩B∣∣A∪B∣J(A,B)=\frac{|A\cap B|}{|A\cup B|}J(A,B)=∣A∪B∣∣A∩B∣
The numerator counts what both sets contain, the denominator everything that appears in at least one. With A={AI,Python,GPU}A={\text{AI},\text{Python},\text{GPU}}A={AI,Python,GPU} and B={Python,GPU,Docker}B={\text{Python},\text{GPU},\text{Docker}}B={Python,GPU,Docker}, two elements are shared and four exist in total, so J=2/4=0.5J=2/4=0.5J=2/4=0.5. With two lists of suspected causes for a network fault, {network,DNS,firewall}{\text{network},\text{DNS},\text{firewall}}{network,DNS,firewall} and {RAM,disk,firewall}{\text{RAM},\text{disk},\text{firewall}}{RAM,disk,firewall}, only one is shared out of five, J=0.2J=0.2J=0.2. Its complement, 1−J1-J1−J, is the Jaccard distance, a true metric.
Jaccard measures literal overlap between sets.
Two answers that say the same thing in different words share few words and score low, so it is the wrong tool for meaning. For embeddings the right tool is cosine similarity.
In vision it is called intersection over union (IoU): the area where a predicted box or mask overlaps the true one, divided by the area they cover together. An object detector's prediction usually counts as correct above an IoU of 0.5, so a detected crack is scored on how well it is outlined.
Low overlap is not good or bad by itself. Comparing several answers a model gives to the same question, a low Jaccard can mean useful diversity of hypotheses or plain inconsistency; which one depends on whether the answers are each right, which the index cannot see.
It ignores the elements absent from both sets, which is its strength for sparse data (two documents are not similar because both lack the word turbine) and the reason it suits presence data far better than a match count would.
For huge sets (deduplicating millions of web pages or log files) it is estimated instead of computed exactly: MinHash gives an unbiased estimate from short signatures, which is how near-duplicate detection runs at scale in the cleaning of training corpora.