Decoder. plain-English AI glossary

Toxicity

● Core

Also called Toxicity

Model output that’s hateful, harassing, or abusive — the language you’d never want your product to say.

Think of it like

A parrot that picked up the worst words from a rough bar and now repeats them to customers.

Example

A chatbot, prompted with a slur, mirrors it back in an insult — the kind of moment that lands a screenshot on social media.

How it actually works

Models learn from internet-scale text, which includes plenty of bile, so the raw capacity to produce toxic language is baked in. Safety training and filters suppress it, but adversarial prompts can still coax it out. "Toxicity" is measured with classifiers and benchmarks, though the label is culturally loaded — what counts as toxic varies by context and audience, which makes both detection and evaluation genuinely hard.

For product teams

A direct brand and trust risk; one toxic output can become a headline.

For engineers

Latent in web-scale training data; suppressed via alignment and output classifiers, but recoverable under adversarial prompts.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome