Refusal prompt
A prompt directive that instructs the model to decline certain requests, usually harmful or outside policy.
Think of it like
Telling a bartender "don't serve anyone who's already drunk," they refuse specific cases.
Example
"Do not write code for malware, exploit sensitive APIs without auth, or generate content that sexualizes minors."
How it actually works
Refusal prompts are a key part of alignment. The tension: make them too strict and users get false refusals (over-refusal). Too loose and you miss actual harms. Optimal prompts are specific, give reasons, and offer alternatives.
For product teams
Critical for compliance and brand trust; measurable through refusal rate metrics.
For engineers
Measure false refusal vs true refusal separately. Over-refusal degrades UX and wastes model capacity.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome