Decoder. plain-English AI glossary

Preference Judgment

● Core

When a human says 'I prefer output A over output B,' capturing subjective quality when objective metrics fail.

Think of it like

Like asking 'which pizza tastes better?' when you can't measure taste objectively.

Example

Two models respond to a creative writing prompt. A human judge reads both and says 'Model A's response is more engaging, even though they're both factually correct.'

How it actually works

Preference judgments are used in RLHF and DPO to train models to match human taste. They're subjective and can vary by rater, so you need multiple raters and clear rubrics. The data is expensive to collect but invaluable for alignment. Risk: raters can have blind spots or be biased toward one model.

For product teams

Preference judgments are how you steer model behavior toward what users actually want.

For engineers

Collect preference data with clear instructions; measure inter-rater agreement; use to train reward models.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome