Decoder. plain-English AI glossary

Jailbreak resistance

▲ Rising

Also called Robustness to Jailbreaks

How well a model’s safety training holds when users get creative about tricking it into breaking the rules.

Think of it like

A bouncer who does not fall for "I left my ID in the car" or any of the fifty other lines they hear a night.

Example

A model refuses a harmful request even when it is wrapped in a role-play, translated to another language, or hidden inside a fake "developer mode" prompt.

How it actually works

Jailbreaks exploit the gap between a model’s helpfulness and its rules, using framing, obfuscation, or persona tricks to route around refusals. Resistance comes from safety fine-tuning plus separate classifier layers, but it is an arms race — new jailbreaks appear constantly. Measured as attack success rate over a red-team suite, ideally trending down.

For product teams

A safe-in-the-demo model is not safe if a Reddit thread can talk it out of its guardrails by Friday.

For engineers

Track attack success rate across jailbreak families; harden with safety fine-tuning plus an independent constitutional classifier.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome