Human Eval
Asking actual people to rate model outputs because metrics often miss what matters.
Think of it like
Like tasting food yourself instead of just looking at the recipe.
Example
You generate two summaries and ask raters: 'Is this summary accurate, concise, and helpful?' You tally their votes.
How it actually works
Humans catch nuance and mistake that automated metrics miss. The tradeoff: it's expensive, slow, and subject to rater bias and disagreement. You need good instructions, training, and quality control. Multiple raters per example help, but don't fully solve bias. Best practice is to use human eval on a hold-out set to validate automated metrics, not as your sole evaluation method.
For product teams
Invest in human eval for high-stakes tasks and to calibrate metrics.
For engineers
Hire annotators, establish inter-rater agreement targets, and build data pipelines to log both model outputs and ratings.
Related
- LLM-as-Judge — Faster alternative.
- A/B Test — Direct comparison method.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome