ML//applied ML
Applied ML is the engineering practice of turning a machine learning method into a prediction that changes a decision in a real system (when to stop a pump, which part to reject, whether a drone battery needs a capacity test), and most of its difficulty lies outside the model: in the question, the data, the labels and the upkeep. The usual commission, five years of data from 300 sensors in a historian, a spreadsheet of failures and *have the AI warn us before something breaks*, becomes a project only once problem framing has written down what exactly is predicted, from which data, for which decision.
Applied ML is the engineering practice of turning a machine learning method into a prediction that changes a decision in a real system (when to stop a pump, which part to reject, whether a drone battery needs a capacity test), and most of its difficulty lies outside the model: in the question, the data, the labels and the upkeep. The usual commission, five years of data from 300 sensors in a historian, a spreadsheet of failures and have the AI warn us before something breaks, becomes a project only once problem framing has written down what exactly is predicted, from which data, for which decision.
Its members follow the order of a project. Industrial labels explains why a failure label is what someone wrote, late and only for what was inspected. Operating regime splits a plant's behaviour into start-up, partial load, full load and cleaning before anything is judged. Physics-based features hands the model what the engineer already knows, and is often the step with the largest payoff. Honest evaluation follows, then MLOps once the model runs. The method selection reads from the top down, and the order of the rungs matters more than their orders of magnitude.
1Threshold or physics ruleone or two variables2Linear or logistic regressionauditable baseline3Random forestrobust without tuning4Gradient boostingtables, best accuracy5Nearest neighbours, SVM or Gaussian processfew data, uncertainty6Clustering and anomaly methodsno labels7ARIMA or boosting on lagsforecasting a series8Neural networkimages, sound, raw signal, text
A network must earn its rent. If gradient boosting on 40 well-chosen features reaches 98 %, a 50-million-parameter network has to justify a GPU, more labels, another deployment pipeline and failures that are harder to explain. It earns it on images, sound, raw signal and text, or when a pretrained model already did the hard work; it does not with a few hundred rows of a table, when the model must extrapolate, or when each decision must be certified. The data-type to network map gives each kind of data something to try first: boosting for process tables and for vibration features, persistence and ARIMA for series, isolation forest without labels, PID or MPC for a control policy.
The error of a learned model has five homes: the data (miscalibrated sensors, misaligned clocks), the labels (repair dates instead of failure dates), the optimization, the generalization (overfitting, leakage) and the drift after training (data drift). Downstream, a miscalibrated probability becomes a misplaced threshold (probability calibration).
Classical methods win when the data fit a meaningful table, the decision must be explained and inference runs on a PLC: 500 boosted trees of depth 6 are about 3,000 comparisons, tens of microseconds. The ladder is the flyswatter rule made concrete: climb a rung when the one below fails on data, and keep it as the fallback.