Alignment tax
Also called Alignment Tax / Safety Tax
The capability you sometimes give up when you make a model safer and more aligned.
Think of it like
Speed bumps in a parking lot: safer for everyone, but they slow you down a little.
Example
After heavy safety fine-tuning, a model refuses a few too many harmless requests and loses a bit of sharpness on hard tasks — that gap is the alignment tax.
How it actually works
Aligning a model — refusal training, safety tuning, RLHF toward cautious behavior — can trade away raw capability or usefulness: over-refusal, blander answers, hedging. The "tax" is that cost. Good alignment work aims to shrink it toward zero, and modern methods have narrowed it a lot, but a tension remains between being maximally helpful and reliably safe.
For product teams
The real trade-off behind "why won't it just answer" — safety and helpfulness pull against each other.
For engineers
The capability/usefulness loss incurred by alignment training; minimized but rarely fully eliminated.
Related
- Refusal training — A main contributor to the tax.
- Safety fine-tuning — The training that can levy it.
- Alignment — The overall goal it is a cost of.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome