Refusal training
Also called Refusal Training
Teaching a model to decline requests it should not fulfill, and to do it gracefully.
Think of it like
Training a bartender to firmly but politely cut someone off — knowing when "no" is the right answer.
Example
The model is fine-tuned on examples where the good response to a harmful request is a clear, polite refusal with a brief reason.
How it actually works
Refusal behavior is taught with curated examples and preference data showing when and how to say no. The craft is calibration: refuse genuinely harmful asks while not over-refusing benign ones that merely pattern-match to something risky. Push too hard and you get an annoying, over-cautious model — a visible face of the alignment tax; too soft and safety gaps open up.
For product teams
Directly shapes the "is this model helpful or annoying" line users feel every day.
For engineers
Fine-tune on curated refusal examples/preferences; the key metric is refusing harmful while minimizing over-refusal.
Related
- Safety fine-tuning — The broader safety phase it is part of.
- Alignment tax — The cost of over-refusal.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome