Rubric Grading
Scoring outputs against a predefined checklist of criteria so grading is consistent, not gut-feel.
Think of it like
Like a restaurant inspection form with checkboxes instead of just 'this place is good or bad.'
Example
A code-generation rubric checks: does it run without error, does it solve the problem, is it readable, is it efficient. Each criterion gets a score.
How it actually works
Rubrics make evaluation reproducible and let you track which aspects improve or degrade. They're useful for both LLM judges and human raters. The catch: writing a good rubric takes iteration, and criteria can conflict (brevity vs. detail). When you give the rubric to an LLM judge, it's only as good as the rubric's clarity.
For product teams
Define what quality means for your product, operationalize it as a rubric, and track scores over time.
For engineers
Build rubrics collaboratively with domain experts; test rubric inter-rater agreement before deploying it.
Related
- LLM-as-Judge — Who might use it.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome