Decoder. plain-English AI glossary

Alignment tax

▲ Rising

Also called Alignment Tax / Safety Tax

The capability you sometimes give up when you make a model safer and more aligned.

Think of it like

Speed bumps in a parking lot: safer for everyone, but they slow you down a little.

Example

After heavy safety fine-tuning, a model refuses a few too many harmless requests and loses a bit of sharpness on hard tasks — that gap is the alignment tax.

How it actually works

Aligning a model — refusal training, safety tuning, RLHF toward cautious behavior — can trade away raw capability or usefulness: over-refusal, blander answers, hedging. The "tax" is that cost. Good alignment work aims to shrink it toward zero, and modern methods have narrowed it a lot, but a tension remains between being maximally helpful and reliably safe.

For product teams

The real trade-off behind "why won't it just answer" — safety and helpfulness pull against each other.

For engineers

The capability/usefulness loss incurred by alignment training; minimized but rarely fully eliminated.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome