ML//model//foundation model

A foundation model is a large neural network pretrained on broad, generic data (usually with self-supervision, predicting held-out parts of the data) and then adapted to many specific tasks, and it is how most vision and language work is now done: nobody trains a model from zero to read work orders or spot scratches on a part. The adaptation takes one of three forms, cheapest last: fine-tuning the whole model (or a small correction to it, LoRA) on a few hundred examples, training only a small head on top of its embeddings, or, for language models, simply asking in the prompt (zero-shot learning).


A foundation model is a large neural network pretrained on broad, generic data (usually with self-supervision, predicting held-out parts of the data) and then adapted to many specific tasks, and it is how most vision and language work is now done: nobody trains a model from zero to read work orders or spot scratches on a part. The adaptation takes one of three forms, cheapest last: fine-tuning the whole model (or a small correction to it, LoRA) on a few hundred examples, training only a small head on top of its embeddings, or, for language models, simply asking in the prompt (zero-shot learning).

The economics explain everything. Someone paid for the pretraining, millions of GPU-hours on data no plant could gather, and the adopter pays for a few hundred labelled examples and an afternoon on one GPU. A CNN for weld inspection that once needed tens of thousands of labelled images now starts from a model that already knows edges, textures and shapes, and learns the defect from a few hundred (transfer learning). The name, coined at Stanford in 2021, stresses that many applications are built on one base, so the base's strengths and defects are inherited by all of them.

Maturity differs by domain.

In vision and language foundation models are common industrial practice; the large language model is the best-known one, and CLIP the bridge between images and text. For industrial time series and for robots (models that map camera images and instructions to actions) there are promising candidates with irregular results, and they remain research.

The hidden costs come with the base. Licences that restrict commercial use, plant data that leave for a vendor's servers when the model is behind an API, and versions that change under the application without notice: a classifier built on a prompt can behave differently after the provider updates the model, so the evaluation set has to be rerun on every version.

A foundation model inherits what it was trained on. If its pretraining data contain few images like the plant's (thermal cameras, X-ray of castings), the head on top learns little, and a small model trained on the plant's own data can beat it at a fraction of the inference cost.

Not every foundation model reads text. Scientific foundation models such as AlphaFold and GenCast are pretrained on protein structures or atmospheric states and predict those systems directly, with no language in between.

In control it stays outside the fast loop. A robot foundation model may choose what to grasp; the joint controllers, the safety limits and the stop button remain classical, for the reasons given in learning-based control.