ML//NLP//sentiment analysis

Sentiment analysis is the NLP task of estimating the attitude a text expresses, most often as positive, negative or neutral, sometimes as a score or per aspect (the delivery was late, the product is excellent), and it is used to read at scale what customers, users or operators say: reviews, support tickets, survey comments, posts about a product launch. A support team that receives 3,000 tickets a day uses it to put the angry ones on top; a product team uses it to see which feature people complain about after a release.


Sentiment analysis is the NLP task of estimating the attitude a text expresses, most often as positive, negative or neutral, sometimes as a score or per aspect (the delivery was late, the product is excellent), and it is used to read at scale what customers, users or operators say: reviews, support tickets, survey comments, posts about a product launch. A support team that receives 3,000 tickets a day uses it to put the angry ones on top; a product team uses it to see which feature people complain about after a release.

It is a supervised classification at heart: text in, a label or a score out. Classic systems counted words from a lexicon of positive and negative terms, then learned weights over word features; fine-tuned encoders such as BERT took over, and an instructed LLM now does it from a single sentence of instructions. Each step handled more of what made the task hard: negation (not bad at all), irony, domain words whose polarity flips (unpredictable is good for a thriller, bad for a brake).

A sentiment score is an opinion about an opinion, so it needs calibrating against people.

Before trusting a dashboard of percentages, a few hundred texts labelled by humans tell how often the model agrees with them, and where it disagrees (sarcasm, mixed reviews, technical complaints written calmly).

Aspect-based sentiment, the useful kind for engineering, ties each opinion to the thing it is about: the drone is stable, the battery dies in ten minutes is positive on stability and negative on endurance, and a single overall label throws that away.

It is evaluated like any classifier, with confusion matrix, precision and recall per class, and agreement with human raters, who themselves disagree on a sizeable share of real texts; that agreement is the ceiling a model can reach.

In plants and fleets its use is narrower than in marketing: operator shift notes and maintenance comments are written to be factual, so what matters there is extracting the fault (named entity recognition), with tone at most a weak signal of urgency.

Shifts in language break it quietly: new slang, a product renamed, a translation layer in between. A model scored once and left running drifts like any other model (out-of-distribution).