ML//NLP//named entity recognition

Named entity recognition (NER) is the NLP task of finding the spans of a text that name things of chosen types (people, organizations, places, dates, quantities, or domain types such as equipment tags, part numbers and fault codes) and labelling each span with its type, and it is the usual first step for turning free text into rows of a database. In *Satya Nadella visited Madrid in September*, a general NER model marks a person, a place and a date; in *Pump P-101 tripped on high bearing temperature at 03:40*, a plant model trained for it marks an equipment tag, a failure mode and a time.


Named entity recognition (NER) is the NLP task of finding the spans of a text that name things of chosen types (people, organizations, places, dates, quantities, or domain types such as equipment tags, part numbers and fault codes) and labelling each span with its type, and it is the usual first step for turning free text into rows of a database. In Satya Nadella visited Madrid in September, a general NER model marks a person, a place and a date; in Pump P-101 tripped on high bearing temperature at 03:40, a plant model trained for it marks an equipment tag, a failure mode and a time.

The output is structure: once every work order has its equipment, its symptom and its date pulled out, the maintenance history in the CMMS can be counted, joined with the historian and searched, which is the point of most industrial text projects. NER is also the first half of relation extraction, which then links the entities (drug X inhibits protein Y), and of the knowledge graphs built from those links.

NER is scored span by span, with precision and recall over entities.

An extracted tag counts only when its boundaries and its type are both right, so P-10 for P-101 is a miss and a false alarm at once, which is exactly the error that breaks a join with the asset register.

Classic systems tagged each token as the beginning, inside or outside of an entity and learned the tags with a sequence model; BERT-style encoders fine-tuned for tagging became the default. An instructed LLM can do it from a prompt and a few examples, and it is a natural use of structured output.

A general model knows general types. Equipment tags, chemical names and the plant's own abbreviations need either a fine-tuned model on a few hundred labelled lines or a dictionary of the asset register used alongside the model, and often both.

Synthetic text is a cheap way to test an extractor: generate sentences around known entities, add typos, paraphrases and distracting names, and measure how many come back intact (synthetic data).

An LLM extractor can invent an entity that is not in the text, a failure the old taggers could not have. Checking that every extracted span actually occurs in the source is a one-line guard worth adding.