Decoder. plain-English AI glossary

Human feedback

● Core

Also called Human Feedback

People judging model outputs — usually by comparing them — to teach the model what "good" looks like.

Think of it like

A taste-test panel picking which of two dishes they prefer, over and over, until the chef learns the room.

Example

Annotators are shown two model replies and click the better one; thousands of these comparisons train the reward model.

How it actually works

Human feedback is how fuzzy human values — helpful, honest, harmless — get injected into training. It usually takes the form of pairwise preferences rather than absolute scores, since people are more consistent at comparing than rating. Those preferences train a reward model or feed directly into methods like DPO. Its limits are human ones: bias, fatigue, and disagreement all leak in.

For product teams

The raw material that turns a raw predictor into something people actually want to talk to.

For engineers

Human preference judgments (often pairwise) used to train reward models or drive preference optimization.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome