ML//training//post-training
Post-training is the family of training stages applied to a model after pre-training, which turn a network that continues text into one that follows instructions, matches preferences and behaves within limits. Pre-training gives the model what it knows; post-training decides how that knowledge comes out when someone asks for it. Every member updates the same weights with a different signal, and what separates them is where that signal comes from and what it costs to collect.
Post-training is the family of training stages applied to a model after pre-training, which turn a network that continues text into one that follows instructions, matches preferences and behaves within limits. Pre-training gives the model what it knows; post-training decides how that knowledge comes out when someone asks for it. Every member updates the same weights with a different signal, and what separates them is where that signal comes from and what it costs to collect.
The usual ladder has three rungs, and a project may skip, repeat or reorder them:
1Pre-training2SFT3Preferences (RLHF or DPO)4Distillation, when it pays
SFT shows the model the answer it should give: pairs of input and desired output, trained with the same next-token loss as pre-training. It is the cheapest rung and the one that teaches format, at the price of imitating single targets.
Preference methods show the model which of two answers is better, which needs no ideal answer to exist. In the classic RLHF recipe a reward model learns the preferences and PPO optimizes the model against it under a KL penalty; DPO reaches a similar result straight from the pairs, with no separate reward model and no online loop. GRPO and outcome rewards from verifiable tasks are the route the reasoning models took; RLAIF and Constitutional AI replace the human judges with a model and a written constitution.
Knowledge distillation transfers a trained model's behaviour into a smaller one, and it is often the last step before deployment.
LoRA is orthogonal to the ladder: it is a way to change a model cheaply, and any rung (SFT most often) can run through it.
Post-training rearranges what pre-training learned.
Every rung moves the model toward useful regions of what pre-training already learned; a capability absent from the base model is not created by any amount of preference data. That is why a base model's quality sets the ceiling and post-training sets how close to it the product gets.
The cost rises along the ladder in a way that matters for planning: SFT needs written answers, preferences need human or model comparisons plus, in the PPO route, a reward model and several copies of the network in memory at once, and every rung needs an evaluation that says whether it helped (evaluation). The safety side runs beside it: red teaming finds failures and its findings come back as training data for the next round. What one run of all this costs, as opposed to the whole programme, is the subject of training cost.