Decoder. plain-English AI glossary

Red Teaming

● Core

Also called Red Teaming

Deliberately attacking your own model to find how it breaks before real attackers do.

Think of it like

Hiring burglars to test your house — better they find the unlocked window than the actual thieves.

Example

Before launch, a team spends weeks trying to make the model produce disallowed content, then feeds every success back into safety training.

How it actually works

Red teaming is adversarial testing applied to AI: skilled people (and increasingly automated systems) probe for jailbreaks, harmful outputs, bias, and misuse paths. The goal isn’t a clean report — it’s a pile of failures you can fix. Findings feed back into alignment training and filters. It’s never "done," because capabilities and attacks keep evolving, so mature teams red team continuously rather than once.

For product teams

The pre-launch stress test that surfaces embarrassing failures on your terms instead of in the press.

For engineers

Structured adversarial probing (human and automated) whose findings feed alignment training and filter updates; a continuous, not one-shot, process.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome