mathematics//statistics//selection bias
Selection bias is the systematic error that appears when the data available for analysis are a sample chosen by some process related to the quantity being estimated, and it is the most common way an engineering dataset lies: the data are a record of what someone decided, or was able, to log. The flight logs that reach the server are those of the drones that came back. The failures in the maintenance system are those that someone inspected and wrote down. The lifetimes in the database stop wherever preventive replacement cut them.
Selection bias is the systematic error that appears when the data available for analysis are a sample chosen by some process related to the quantity being estimated, and it is the most common way an engineering dataset lies: the data are a record of what someone decided, or was able, to log. The flight logs that reach the server are those of the drones that came back. The failures in the maintenance system are those that someone inspected and wrote down. The lifetimes in the database stop wherever preventive replacement cut them.
Each of these skews the answer in a predictable direction. Logs from returning drones understate the conditions that bring drones down, because the worst flights left no log (survivorship bias). A fleet that replaces bearings at 8,000 hours never sees one fail at 12,000; treating its records as complete lives underestimates durability, and the remedy is to treat the replaced units as censored data, whose lives are known only to exceed their age. Labels written only for inspected machines say nothing about the ones nobody opened (industrial labels).
More data does not repair selection bias.
It shrinks the random error around a value that is still wrong, so a biased estimate becomes more confidently biased (estimation). The only defence is to ask how each row reached the table, and to model or remove the selection.
Fitting only the failures is the textbook case in reliability. With five bearings failed and fifteen still running at 1,200 h, a Weibull fit on the five alone gives a characteristic life of about 860 h instead of about 2,100 h: bearings would be changed at mid-life with full statistical confidence (Weibull distribution).
Machine learning meets it as a training set that differs from deployment: data from the operating regimes someone thought to log, or from before a controller change, describe a different population from the one the model will see.
A closed loop is a selection mechanism of its own. A controller chooses its inputs from the state, so logged data cover only what the loop allowed, and the relation between input and output in them is shaped by the controller (causal inference).
It is distinct from leakage, where information from the answer slips into the inputs (data leakage); both produce results that look better than the truth, and both are found by asking where each value came from.