Constitutional AI
Also called Constitutional AI
Aligning a model against a written set of principles instead of case-by-case human labels.
Think of it like
Like giving an employee a clear code of conduct and letting them self-correct against it, rather than a manager approving every single action.
Example
A model critiques and revises its own answers against rules like “be helpful but don’t assist harm,” and those revisions train the final model.
How it actually works
Constitutional AI, introduced by Anthropic, writes down explicit principles (a “constitution”) and has the model use them to critique and improve its own outputs, generating training data with minimal human labeling. It makes the values steering the model explicit and auditable, and scales better than hand-labeling every case. It still depends on the wisdom of the principles and the model’s ability to apply them faithfully.
For product teams
Values you can read and revise, not buried in opaque labels — useful when you need to explain why a model behaves as it does.
For engineers
Self-critique/revision against a written principle set to generate alignment data; a structured form of AI feedback.
Related
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome