Decoder. plain-English AI glossary

Human Eval

● Core

Asking actual people to rate model outputs because metrics often miss what matters.

Think of it like

Like tasting food yourself instead of just looking at the recipe.

Example

You generate two summaries and ask raters: 'Is this summary accurate, concise, and helpful?' You tally their votes.

How it actually works

Humans catch nuance and mistake that automated metrics miss. The tradeoff: it's expensive, slow, and subject to rater bias and disagreement. You need good instructions, training, and quality control. Multiple raters per example help, but don't fully solve bias. Best practice is to use human eval on a hold-out set to validate automated metrics, not as your sole evaluation method.

For product teams

Invest in human eval for high-stakes tasks and to calibrate metrics.

For engineers

Hire annotators, establish inter-rater agreement targets, and build data pipelines to log both model outputs and ratings.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome