Moderation
Also called Moderation
The overall process of screening content against a policy and deciding what to allow, block, or flag.
Think of it like
An editor with a house style guide — deciding what runs, what gets cut, and what needs a second look.
Example
A platform routes every user message and model reply through a moderation step that scores for hate, self-harm, and sexual content, then applies the policy.
How it actually works
Moderation is the umbrella practice; content filters and classifiers are its tools. It combines automated scoring with policy definitions and, often, human review for edge cases. The hard part isn’t the classifier — it’s the policy: drawing lines that hold across cultures, contexts, and gray areas, and handling the appeals when the automated call is wrong. It’s as much governance as engineering.
For product teams
Where your values become rules — the policy decisions here define what your product will and won’t say.
For engineers
Policy-driven screening combining classifiers, thresholds, and human review; the classifier is easy, the policy and edge cases are hard.
Related
- Content Filter — The classifier tool it relies on.
- Toxicity — A primary category it screens.
- Harmful Content — Handling of harmful material broadly.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome