Red Teaming
Also called Red Teaming
Deliberately attacking your own model to find how it breaks before real attackers do.
Think of it like
Hiring burglars to test your house — better they find the unlocked window than the actual thieves.
Example
Before launch, a team spends weeks trying to make the model produce disallowed content, then feeds every success back into safety training.
How it actually works
Red teaming is adversarial testing applied to AI: skilled people (and increasingly automated systems) probe for jailbreaks, harmful outputs, bias, and misuse paths. The goal isn’t a clean report — it’s a pile of failures you can fix. Findings feed back into alignment training and filters. It’s never "done," because capabilities and attacks keep evolving, so mature teams red team continuously rather than once.
For product teams
The pre-launch stress test that surfaces embarrassing failures on your terms instead of in the press.
For engineers
Structured adversarial probing (human and automated) whose findings feed alignment training and filter updates; a continuous, not one-shot, process.
Related
- Jailbreak — The attacks it hunts for.
- Jailbreak Prompt — It maintains libraries of these.
- AI Safety — Feeds the broader safety effort.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome