Decoder. plain-English AI glossary

Jailbreak Prompt

▲ Rising

Also called Jailbreak Prompt

A specific, often reusable script that reliably talks a model past its safety training.

Think of it like

A lock-picking template passed around online — the exact wiggle that pops a particular model.

Example

The old "DAN" ("Do Anything Now") prompt told the model to play an unrestricted alter ego, and got copied and remixed across forums until it was patched.

How it actually works

Where "jailbreak" is the general act, a jailbreak prompt is a concrete artifact — a piece of text that packages the trick. They spread and mutate: someone finds a framing that works, it circulates, the provider trains against it, a variant appears. Studying them is how safety teams find gaps, which is why red teams build libraries of known jailbreak prompts to test each new model against.

For product teams

These circulate publicly, so treat them as known attacks to test against before shipping, not surprises.

For engineers

Concrete reusable attack strings; providers patch them and adversaries iterate, so red teams maintain regression suites of them.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome