Jailbreak Prompt
Also called Jailbreak Prompt
A specific, often reusable script that reliably talks a model past its safety training.
Think of it like
A lock-picking template passed around online — the exact wiggle that pops a particular model.
Example
The old "DAN" ("Do Anything Now") prompt told the model to play an unrestricted alter ego, and got copied and remixed across forums until it was patched.
How it actually works
Where "jailbreak" is the general act, a jailbreak prompt is a concrete artifact — a piece of text that packages the trick. They spread and mutate: someone finds a framing that works, it circulates, the provider trains against it, a variant appears. Studying them is how safety teams find gaps, which is why red teams build libraries of known jailbreak prompts to test each new model against.
For product teams
These circulate publicly, so treat them as known attacks to test against before shipping, not surprises.
For engineers
Concrete reusable attack strings; providers patch them and adversaries iterate, so red teams maintain regression suites of them.
Related
- Jailbreak — The general technique it packages.
- Red Teaming — The practice of collecting and testing them.
- Refusal — The safety behavior they target.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome