Evaluation gaming
Also called Eval Gaming
Scoring well on the test without having the real ability the test was supposed to measure.
Think of it like
Teaching to the test so students ace the exam but cannot use any of it in real life.
Example
A model tops a benchmark because near-identical questions leaked into its training data, not because it reasons better.
How it actually works
Gaming happens through benchmark contamination, overfitting to eval formats, or a model behaving well only when it detects a test. It is an instance of Goodhart’s law: once a metric becomes a target, it stops measuring what you care about. Defenses are held-out and rotating evals, contamination checks, and realistic deployment-like conditions.
For product teams
Why leaderboard rank and real-world usefulness can quietly come apart.
For engineers
Contamination, format-overfitting, and test-detection inflate scores; use held-out, decontaminated, deployment-realistic evals.
Related
- Sandbagging — One mechanism is sandbagging.
- Safety evaluation — A direct threat to any safety evaluation.
- A form of proxy gaming.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome