Decoder. plain-English AI glossary

Refusal prompt

▲ Rising

A prompt directive that instructs the model to decline certain requests, usually harmful or outside policy.

Think of it like

Telling a bartender "don't serve anyone who's already drunk," they refuse specific cases.

Example

"Do not write code for malware, exploit sensitive APIs without auth, or generate content that sexualizes minors."

How it actually works

Refusal prompts are a key part of alignment. The tension: make them too strict and users get false refusals (over-refusal). Too loose and you miss actual harms. Optimal prompts are specific, give reasons, and offer alternatives.

For product teams

Critical for compliance and brand trust; measurable through refusal rate metrics.

For engineers

Measure false refusal vs true refusal separately. Over-refusal degrades UX and wastes model capacity.

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome