Jailbreak resistance
Also called Robustness to Jailbreaks
How well a model’s safety training holds when users get creative about tricking it into breaking the rules.
Think of it like
A bouncer who does not fall for "I left my ID in the car" or any of the fifty other lines they hear a night.
Example
A model refuses a harmful request even when it is wrapped in a role-play, translated to another language, or hidden inside a fake "developer mode" prompt.
How it actually works
Jailbreaks exploit the gap between a model’s helpfulness and its rules, using framing, obfuscation, or persona tricks to route around refusals. Resistance comes from safety fine-tuning plus separate classifier layers, but it is an arms race — new jailbreaks appear constantly. Measured as attack success rate over a red-team suite, ideally trending down.
For product teams
A safe-in-the-demo model is not safe if a Reddit thread can talk it out of its guardrails by Friday.
For engineers
Track attack success rate across jailbreak families; harden with safety fine-tuning plus an independent constitutional classifier.
Related
- Jailbreak — Resisting the technique known as a jailbreak.
- Constitutional classifier — One defensive layer is a constitutional classifier.
- Refusal rate — Measured partly by refusal rate on genuinely harmful asks.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome