Decoder. plain-English AI glossary

Refusal

● Core

Also called Refusal

When a model declines to answer because the request crosses a safety line.

Think of it like

A pharmacist who won’t fill a prescription that looks forged, no matter how politely you ask.

Example

Asked for step-by-step instructions to synthesize a dangerous toxin, the model declines and offers to explain the chemistry at a safe, general level instead.

How it actually works

Refusal is a trained behavior: models learn to recognize categories of harmful requests and respond with a decline rather than comply. Done well, it’s the visible edge of alignment. The craft is calibration — refusing genuinely harmful asks while still helping with the vast majority of legitimate ones. Refuse too eagerly and you get over-refusal; refuse too little and safety training failed.

For product teams

The user-facing face of your safety policy; its tone and accuracy shape how trustworthy the product feels.

For engineers

A learned response to harmful-request classes; quality is the calibration between correct declines and false positives.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome