ML//model//foundation model//scientific foundation model

A scientific foundation model is a large neural network trained on the data of a physical, chemical or biological system (protein sequences and structures, decades of atmospheric reanalysis, molecules, materials) to predict that system's behaviour, and it is how machine learning moved from documents and conversations to the systems science studies. Its inputs and outputs are the system's own representation, never text: amino-acid sequences in and atomic coordinates out for AlphaFold, the state of the atmosphere on a global grid in and its future states out for GenCast. Most of them are not language models at all, even when they share the transformer and the diffusion machinery.


A scientific foundation model is a large neural network trained on the data of a physical, chemical or biological system (protein sequences and structures, decades of atmospheric reanalysis, molecules, materials) to predict that system's behaviour, and it is how machine learning moved from documents and conversations to the systems science studies. Its inputs and outputs are the system's own representation, never text: amino-acid sequences in and atomic coordinates out for AlphaFold, the state of the atmosphere on a global grid in and its future states out for GenCast. Most of them are not language models at all, even when they share the transformer and the diffusion machinery.

Two examples carry the idea. AlphaFold (DeepMind) predicts the three-dimensional structure of a protein from its sequence, a problem that took months of laboratory work per protein; its 2020 version reached accuracy comparable to experiment on many targets, and the work shared the 2024 Nobel Prize in Chemistry. GenCast (DeepMind, December 2024) produces an ensemble of 50 or more possible 15-day global weather trajectories at 0.25° resolution, each in about eight minutes on one accelerator, and in the published evaluation it beat the European centre's operational ensemble on 97 % of the variable and lead-time targets tested. A forecast that took hours on a supercomputer became a few minutes of inference, which changes how many scenarios can be run.

These models learn the system from its data, so they inherit the data's reach.

A weather model trained on past decades has seen few events like the most extreme ones a warming climate produces, and a structure predictor is weaker on proteins unlike anything in its training set; physics-based models and measurements stay as the check.

They are often emulators: trained to reproduce, far faster, what a physical simulation or an experiment would give. That makes them simulators with a learned core, useful for running thousands of what-ifs, and bounded by how well the training data cover the case at hand.

Probabilistic output matters as much as accuracy. GenCast is useful because its ensemble gives the spread of possible weather, which is what a grid operator planning wind output or a drone operator planning a mission needs (forecast calibration).

The idea of combining several of them (a weather model feeding an energy-demand model feeding a grid model) is appealing and largely untested; errors and uncertainties compound across the chain.

For a company the general foundation model lessons apply: someone else paid for the pretraining, the adopter brings a few domain examples, and the model's blind spots become the application's.