Red teaming
Deliberately trying to break an AI — finding the prompts that make it say or do things it shouldn't.
Think of it like
Hiring a burglar to test your locks before a real thief shows up.
Example
A team spends a week trying to get the chatbot to leak its system prompt, produce harmful content, or ignore its guardrails.
How it actually works
Red teamers probe for jailbreaks, harmful outputs, bias, privacy leaks, and edge-case failures — anything the model should refuse but might not. It's both a pre-launch audit and an ongoing practice, because new attack techniques keep emerging.
For product teams
The security audit for AI — find the failures before users do.
For engineers
Adversarial probing for jailbreaks, harmful outputs, bias, and prompt-injection vectors.
Related
- Tests the guardrails.
- Evals — A form of eval focused on safety.
- Prompt injection — Tries to exploit prompt injection.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome