ML//applied ML//problem framing

Problem framing is the step of an ML project that writes down, before any model exists, exactly what will be predicted, from which data available at decision time, to change which decision, at what cost per error and against which baseline, and it decides whether the project can succeed long before the algorithm does. A target written as *the state of the machine* is no label; *failure of the drive-end bearing within the next 7 days* is one.


Problem framing is the step of an ML project that writes down, before any model exists, exactly what will be predicted, from which data available at decision time, to change which decision, at what cost per error and against which baseline, and it decides whether the project can succeed long before the algorithm does. A target written as the state of the machine is no label; failure of the drive-end bearing within the next 7 days is one.

Each of the five answers kills a common failure. The target must be a label someone can actually produce. The data must exist when the decision is taken: an oil analysis that takes ten days to come back is no input for today. The decision must change something (stop, order a spare, slow down); if nothing changes, the project is an expensive dashboard. And the cost of each error sets the threshold later: in a fleet of delivery drones a false alarm costs half an hour of inspection and a missed failure costs a drone on the ground, and that asymmetry fixes where to cut (decision threshold).

The most expensive error lives in no equation: it is the ill-posed question.

Predicting what nobody needs to decide, estimating what is not observable, or optimizing the wrong cost produces results that are precise and exactly useless, and every layer built on such a question inherits it.

The baseline is the veteran's rule written down and measured: if vibration exceeds 7 mm/s, inspect. The model has to beat it with a margin that pays for its own upkeep (retraining, monitoring, a pipeline that computes its features the same way and on time), and when it does not, the rule ships. A baseline also exposes leakage: a first model that beats the expert by a mile is usually copying the answer.

The framing fixes the evaluation. A question about a machine the model has never seen demands a group split, one about next month a temporal one (train-validation-test split), and the cost of each error picks the metric (precision and recall).

The value lies in the decision a prediction changes, which is why predictive maintenance starts from the maintenance plan it will alter. The next steps are the labels (industrial labels) and the features (physics-based features).