Preference data
Human (or AI) judgments of “this answer is better than that one,” used to teach models taste.
Think of it like
Like A/B taste tests: you don’t score each dish alone, you just say which of two you’d rather eat.
Example
A labeler sees two model replies to the same question and marks the clearer, more helpful one as preferred.
How it actually works
Rather than absolute scores, preference data captures relative rankings between responses, which humans give far more reliably than numeric grades. This pairwise signal trains reward models and drives methods like DPO. Its quality is everything: inconsistent, rushed, or biased comparisons teach the model the wrong preferences with perfect efficiency.
For product teams
The raw material of alignment — the model’s manners are only as good as these comparisons.
For engineers
Pairwise/ranked response comparisons; the training signal for reward models and preference-optimization losses.
Related
- Reward model — Trains a reward model.
- DPO — Consumed directly by DPO.
- HITL — Gathered via HITL.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome