Toxicity
Also called Toxicity
Model output that’s hateful, harassing, or abusive — the language you’d never want your product to say.
Think of it like
A parrot that picked up the worst words from a rough bar and now repeats them to customers.
Example
A chatbot, prompted with a slur, mirrors it back in an insult — the kind of moment that lands a screenshot on social media.
How it actually works
Models learn from internet-scale text, which includes plenty of bile, so the raw capacity to produce toxic language is baked in. Safety training and filters suppress it, but adversarial prompts can still coax it out. "Toxicity" is measured with classifiers and benchmarks, though the label is culturally loaded — what counts as toxic varies by context and audience, which makes both detection and evaluation genuinely hard.
For product teams
A direct brand and trust risk; one toxic output can become a headline.
For engineers
Latent in web-scale training data; suppressed via alignment and output classifiers, but recoverable under adversarial prompts.
Related
- Content Filter — The classifier layer that screens it.
- Moderation — The broader screening process.
- Jailbreak — Attacks that try to elicit it.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome