Harmful Content
Also called Harmful Content
The categories of output a model shouldn’t produce because they can cause real damage.
Think of it like
The list of things a responsible publisher won’t print — bomb recipes, targeted harassment, exploitation — no matter who asks.
Example
A safety policy enumerates harmful categories — weapons synthesis, child exploitation, credible violence, self-harm encouragement — that the model must always refuse.
How it actually works
"Harmful content" is the catch-all for what safety systems are built to prevent. Some categories are near-universal and severe (weapons of mass destruction, child sexual abuse material), where refusal is absolute; others are context-dependent and debated, where reasonable policies differ. Defining the taxonomy is the foundational act of safety work — every filter, classifier, and refusal is downstream of where you draw these lines.
For product teams
The policy taxonomy that everything else in your safety stack is built to enforce.
For engineers
The harm taxonomy underlying refusals and classifiers; some categories are absolute, others context-dependent and contested at the policy layer.
Related
- Refusal — The behavior triggered when it’s detected.
- Safety Classifier — The tool that scores for it.
- Moderation — The process that screens it.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome