Content policy
Also called Content Policy
The written rules for what content a system will and will not produce or allow.
Think of it like
A restaurant’s posted menu of what they do and do not serve — the line the kitchen will not cross.
Example
A model’s content policy forbids weapons instructions and explicit content, and that document is what the safety classifiers are trained to enforce.
How it actually works
Content policy defines the categories — violence, sexual content, self-harm, illegal acts — and the thresholds for each, translating fuzzy values into enforceable rules. It is the source of truth that both human reviewers and automated classifiers point back to. Drawing the lines is genuinely hard, since context changes whether the same words are fine or harmful.
For product teams
The spec that decides what your product refuses — get it wrong in either direction and you feel it.
For engineers
The labeled rule set that grounds classifier training and human review; ambiguity here becomes inconsistency downstream.
Related
- Trust and safety — Owned and enforced by trust and safety.
- Usage policy — Sibling document governing user conduct is the usage policy.
- CSAM detection — Its strictest category is CSAM detection.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome