Lexicon · Capability & training

Post-training

Plain English. Everything done to a model after pretraining: supervised fine-tuning on curated examples, reinforcement learning from feedback or verifiable rewards, safety training. Pretraining produces a raw text predictor; post-training turns it into a product that follows instructions, reasons at length and declines to help with bioweapons.

Why it moves money. The differentiation between frontier labs has migrated here. Pretraining compute is increasingly available to anyone with capital; the post-training recipe — what data, what rewards, in what order — is the part labs guard. It is also far cheaper than pretraining, which means capability gaps can close faster than capex comparisons suggest: a rival doesn't need your cluster to copy your behaviour, only a good enough recipe.

What to watch. How much of a lab's compute and headcount shifts from pretraining to post-training, and how quickly open releases replicate closed-model behaviour. When the recipe leaks or gets published, the gap it protected closes on a timescale of months.

From the signals. Kimi K3 shipped the whole training recipe, not just the weights — recipe disclosure as a competitive act. GLM-5.3 shipped open with its makers calling the cyber capability gain emergent, a reminder that post-training doesn't fully control what surfaces.

← All terms