ML//NLP
NLP (natural language processing) is the field of computing that makes programs read, classify, extract from and generate human language, and it is what turns the text a business or a plant produces (work orders, operator logs, emails, manuals, reviews) into something a system can act on. For decades it was a collection of separate tasks, each with its own trained model: named entity recognition to pull out names, places and part numbers, sentiment analysis to grade opinions, translation, summarization, question answering.
NLP (natural language processing) is the field of computing that makes programs read, classify, extract from and generate human language, and it is what turns the text a business or a plant produces (work orders, operator logs, emails, manuals, reviews) into something a system can act on. For decades it was a collection of separate tasks, each with its own trained model: named entity recognition to pull out names, places and part numbers, sentiment analysis to grade opinions, translation, summarization, question answering.
The large language model changed the shape of the field more than its goals. One instructed model now does most of those tasks from a sentence of instructions and a few examples (zero-shot learning), so a team that once trained and maintained five models can start with one prompt. The tasks did not disappear: they became the way an LLM's output is specified and evaluated, and their old metrics (precision and recall, BLEU, ROUGE) still score it.
The general model wins on breadth; the specialist still wins on cost, latency, privacy and narrow precision.
A small fine-tuned extractor that runs on the plant's own server in a few milliseconds, never sends a work order outside, and gets the plant's part-number format right can beat a frontier LLM on the one job it was built for.
The classic pipeline went from raw text to tokens, then to features (one-hot counts, TF-IDF, later Word2Vec vectors), then to a task model such as logistic regression or a sequence tagger. The modern one goes from text to a tokenizer and a pretrained transformer, adapted by a prompt or by fine-tuning.
BERT marks the middle step: one pretrained encoder, fine-tuned per task, the default for classification and extraction from 2019 until instructed LLMs took over generation and much of the rest.
Industrial text is hard text: abbreviations, shorthand, mixed languages, tag names and misspellings typed on a tablet with gloves on. A model that reads newspapers well can still misread P-101 trip, brg temp hi; checking it on a sample of the plant's own text is the first step of any project.
Choosing between an LLM call and a specialist is an engineering decision with numbers: volume of documents, cost per call, latency budget, where the data may go, and the accuracy measured on a labelled sample, not on a public benchmark.