mathematics//statistics//case-control study
A case-control study is an observational design that starts from the outcome: it gathers units that already have a condition (the cases) and comparable units that do not (the controls), and compares how often each group was exposed to a suspected cause. It is the design of choice when the outcome is rare or slow, because waiting for enough outcomes to happen in a followed group would take decades or a fleet of thousands. The famous example is Doll and Hill's 1950 study of lung cancer patients in London hospitals against matched patients without it: heavy smokers were enormously over-represented among the cases, a first strong sign that smoking causes the disease, later confirmed by following doctors forward in time.
A case-control study is an observational design that starts from the outcome: it gathers units that already have a condition (the cases) and comparable units that do not (the controls), and compares how often each group was exposed to a suspected cause. It is the design of choice when the outcome is rare or slow, because waiting for enough outcomes to happen in a followed group would take decades or a fleet of thousands. The famous example is Doll and Hill's 1950 study of lung cancer patients in London hospitals against matched patients without it: heavy smokers were enormously over-represented among the cases, a first strong sign that smoking causes the disease, later confirmed by following doctors forward in time.
The quantity it can estimate is the odds ratio. With aaa exposed and ccc unexposed cases, bbb exposed and ddd unexposed controls,
OR=a/cb/d=adbc\mathrm{OR} = \frac{a/c}{b/d} = \frac{ad}{bc}OR=b/da/c=bcad
and because the analyst chose how many cases and controls to collect, the design cannot give the risk of the outcome itself; when the outcome is rare, the odds ratio approximates the relative risk (Cornfield's argument). The same logic serves an engineer investigating forty failed pumps against a hundred and twenty healthy ones of the same model: compare supplier lot, duty cycle and installation crew between the two groups instead of waiting years for new failures.
A case-control study compares groups and speaks about groups.
A difference in means between patients and controls says that the groups differ on average; it does not say that any one patient carries the difference, which is the limit normative modelling was built to get past in brain imaging.
Its main weakness is the choice of controls. If controls are not drawn from the population that produced the cases, the comparison measures that difference instead of the exposure (selection bias); looking back also invites recall bias, cases searching their memory harder than controls.
It shows association. Confounders that travel with the exposure produce the same pattern, so causal claims need the reasoning of causal inference or a design that follows exposure forward.
Small group differences found in small samples are often false positives, the classic failure of early brain-imaging studies of psychiatric disorders, whose effects shrank when pooled across thousands of scans (type I and type II errors).
In maintenance and reliability it hides inside many root-cause reviews: a list of failed units compared with a list of survivors is a case-control study, and inherits all its biases.