Alignment
The work of making a model’s behavior match human intentions and values, not just its literal training objective.
Think of it like
Like the gap between a genie granting your exact words and one granting what you actually meant — alignment is aiming for the second.
Example
A model that refuses to help build a weapon but happily explains chemistry safely is showing alignment between capability and intent.
How it actually works
Alignment spans everything from surface manners to deep safety: the model should be helpful, honest, and harmless, and should pursue what we want rather than a proxy that merely correlates with it. Techniques like RLHF, DPO, and constitutional methods are practical alignment tools. The hard part is that we often can’t fully specify our own values, and optimizing a proxy can drift in unintended ways — the core worry as models get more capable.
For product teams
The difference between a model that’s smart and one that’s trustworthy; it’s a product requirement, not an afterthought.
For engineers
Making learned behavior track intended goals/values rather than a mis-specified proxy objective.
Related
- RLHF — Pursued via RLHF.
- Constitutional AI — Also pursued via Constitutional AI.
- Reward model — A defense against reward hacking in the reward model.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome